
< session />
Thu, April 29Reliability, Observability & SecurityAI-Native Software
Production AI agent systems can appear healthy by traditional operational measures while delivering a steadily degrading user experience. Latency remains low, tools execute successfully, databases remain healthy, and dashboards stay green, yet responses become increasingly vague, hesitant, and less useful. The problem is not operational failure, but behavioural degradation.
In this session, you will examine five production incidents from LangGraph-based agent systems that expose the gap between infrastructure health and user experience. You will see how unbounded memory injection reduced response quality without triggering alerts, how semantic caching served confidently incorrect answers, why OAuth refresh races emerged under concurrent streaming workloads, how conventional circuit-breaker strategies increased cloud costs, and how observability broke down across Server-Sent Events. Working from production artifacts including logs, quality metrics, traces, latency distributions, authentication failures, dashboards, and code, you will explore the resulting patterns for memory lifecycle management, cache correctness, streaming-safe authentication, latency-aware failure handling, and observability across protocol boundaries.
What You Will Learn
How behavioural degradation can emerge even when traditional operational metrics remain healthy
Production patterns for memory management, cache correctness, authentication, failure handling, and observability
Practical lessons drawn from real incidents, investigations, and remediation efforts in production agent systems
Who Should Attend
AI Engineers, Platform Engineers, SREs and Reliability Engineers, Staff and Principal Engineers, Software Architects, Technical Leads, and teams operating LangGraph, CrewAI, AutoGen, or custom agent platforms.
< speaker_info />
Tuhin Sharma is Senior Principal Data Scientist at Redhat in the Data & AI team. Prior to that, he worked at Hypersonix as an AI architect and at IBM Watson as Data Scientist. He also co-founded and has been CEO of Binaize (backed by Techstars), a website conversion intelligence product for e-commerce SMBs. He received a master's degree from IIT Roorkee and a bachelor's degree from IIEST Shibpur in Computer Science. He loves to code and collaborate on open-source projects. He is one of the top 20 contributors of pandas. He has 4 research papers and 5 patents in the fields of AI and NLP. He is a reviewer of the IEEE MASS conference, Springer nature and Packt publication in the AI track. He writes deep learning articles for O'Reilly in collaboration with the AWS MXNET team. He is a regular speaker at prominent AI conferences like O'Reilly Strata & AI, PyCon, PyData, ODSC, GIDS, Devconf, Datahack Summit etc.
Soham is a Principal Engineer at Red Hat, where he builds agentic AI systems that survive the jump from demo to production. Based in Bangalore, he owns the full stack, from React front-ends through vLLM serving on Kubernetes, with reliability, observability, and security built in from day one.
Soham's production work, from enabling AI for multi-tenant data infrastructure to open-source agentic templates now running across Red Hat's Data & AI organization, reflects a rare cross-domain fluency across backend, security, and AI/ML. A published author and speaker, Soham cuts through the hype on what it actually takes to run AI at scale: closing the demo-to-production gap, taming cascading failures in multi-agent systems, and the engineering discipline that turns fragile prototypes into systems you can trust.