Mahamood Hussain Mirza

Mahamood Hussain Mirza

Mahamood Hussain Mirza

Software Engineer, Biztegy Analytics INC.

Title of Talk:

AI-Driven Reliability Engineering: Transforming DevOps and SRE Practices for Resilient Cloud Systems


Abstract

In today's hyper-scale cloud environments, traditional DevOps and Site Reliability Engineering (SRE) practices face unprecedented challenges in maintaining availability, scalability, and resilience amid increasing complexity and dynamic workloads. This keynote explores how Artificial Intelligence (AI) and Machine Learning (ML) are revolutionizing reliability engineering by shifting from reactive incident response to proactive, intelligent automation. Drawing on over 10 years of hands-on experience in AWS, GCP, Kubernetes, Terraform, and CI/CD pipelines, I will present real-world case studies on implementing AIOps for anomaly detection, predictive maintenance, root cause analysis, and automated remediation. The talk will also highlight hybrid AI approaches, including Proximal Policy Optimization (PPO) combined with generative models for innovative system design, and interpretable fuzzy decision support systems (FDSS) for monitoring and early intervention in complex operational data. Attendees will gain actionable insights into key challenges such as alert fatigue reduction, SLO management, chaos engineering with AI, and ethical considerations in autonomous systems. By bridging cutting-edge research with industry practices, this presentation demonstrates how AI empowers SREs to achieve higher reliability, lower MTTR, and sustainable cloud operations-paving the way for the next generation of intelligent infrastructure


Brief Profile
I am a seasoned DevOps and Site Reliability Engineer (SRE) with over a decade of experience architecting resilient cloud infrastructure at scale. Currently serving as a DevOps/SRE at Gainwell Technologies, I have held key roles at organizations including Charles Schwab, Nordstrom, and NBC Universal. My expertise spans AWS, GCP, Kubernetes, Docker, Terraform, Jenkins, and advanced monitoring solutions, with a strong focus on automation, incident management, and chaos engineering. In parallel, I am an active AI/ML researcher specializing in reinforcement learning, generative models, and interpretable decision systems. My recent work includes hybrid PPO-ESGAN frameworks for innovative design and fuzzy decision support systems for academic/performance monitoring, with publications in relevant venues. As a passionate advocate for AI-augmented reliability engineering, I frequently share insights on bridging research and practice. I hold an M.S. in Computer Information Systems from New England College and a B.Tech. in Computer Science & Engineering. I am an IEEE member and am pursuing BCS Fellowship. This keynote reflects my commitment to advancing intelligent, reliable systems that drive operational excellence in the era of AI.