
< session />
Thu, April 29AI-Native SoftwarePlatform Engineering & DevOpsReliability, Observability & Security
Most AI agent tutorials focus on capabilities, but production systems often encounter a different set of problems. Agents that perform well in development can behave differently when exposed to real users, long-lived sessions, concurrent workloads, streaming interactions, and operational requirements. Authentication can behave differently under load, observability can break across protocol boundaries, memory can accumulate over time, and resiliency mechanisms designed for traditional services can produce unexpected outcomes in LLM-powered systems.
In this live session, you will see a production-ready agent platform built from an empty directory into an authenticated, observable, streaming application, with no prerecorded demonstrations or hidden infrastructure. As the architecture evolves, you will examine the production incidents behind each design decision, including OAuth refresh races under concurrent streaming workloads, distributed traces disappearing at Server-Sent Event boundaries, uncontrolled memory growth affecting response quality, and the behaviour of retries, caching, and circuit breakers in agent-based systems. The implementation brings together an MCP server, a LangGraph agent, a backend-for-frontend layer, OpenTelemetry tracing, checkpointing, configuration management, and streaming user interactions.
What You Will Learn
How to assemble a production-ready agent platform with authentication, streaming, tracing, and operational controls
How production failures can emerge around memory growth, authentication, observability, caching, and resiliency mechanisms
How architectural patterns can help agent systems remain observable, maintainable, and reliable after deployment
Who Should Attend
AI Engineers, Platform Engineers, Full-Stack Engineers building agent applications, SREs and Reliability Engineers, Technical Leads, Software Architects, and teams deploying LangGraph or MCP-based agent systems.
< speaker_info />
Tuhin Sharma is Senior Principal Data Scientist at Redhat in the Data & AI team. Prior to that, he worked at Hypersonix as an AI architect and at IBM Watson as Data Scientist. He also co-founded and has been CEO of Binaize (backed by Techstars), a website conversion intelligence product for e-commerce SMBs. He received a master's degree from IIT Roorkee and a bachelor's degree from IIEST Shibpur in Computer Science. He loves to code and collaborate on open-source projects. He is one of the top 20 contributors of pandas. He has 4 research papers and 5 patents in the fields of AI and NLP. He is a reviewer of the IEEE MASS conference, Springer nature and Packt publication in the AI track. He writes deep learning articles for O'Reilly in collaboration with the AWS MXNET team. He is a regular speaker at prominent AI conferences like O'Reilly Strata & AI, PyCon, PyData, ODSC, GIDS, Devconf, Datahack Summit etc.
Soham is a Principal Engineer at Red Hat, where he builds agentic AI systems that survive the jump from demo to production. Based in Bangalore, he owns the full stack, from React front-ends through vLLM serving on Kubernetes, with reliability, observability, and security built in from day one.
Soham's production work, from enabling AI for multi-tenant data infrastructure to open-source agentic templates now running across Red Hat's Data & AI organization, reflects a rare cross-domain fluency across backend, security, and AI/ML. A published author and speaker, Soham cuts through the hype on what it actually takes to run AI at scale: closing the demo-to-production gap, taming cascading failures in multi-agent systems, and the engineering discipline that turns fragile prototypes into systems you can trust.