ALL SPEAKERS
AT A GLANCE
Role
Product Manager
Organisation
BSH Hausgeräte GmbH
Category
Community
ABOUT THE SPEAKER
András Bocsák is a Product Manager at BSH Hausgeräte GmbH with a focus on conversational AI and production-grade AI systems. Over the past nine years, he has led digital product initiatives across the Home Connect ecosystem. Today, he works on AI Assist, the customer-facing AI assistant in the Home Connect app, where he helps define testing and evaluation strategies for large language model applications. His areas of interest include AI quality assurance, automated evaluations, observability, LLM judges, and continuous improvement of customer-facing AI systems. He is passionate about turning emerging technologies into products that deliver practical value at scale.
TALK
Testing the Untestable? Quality Control for Multi-Agent LLM Systems in Production
“It works on my prompt" is not a QA strategy, but what else can you do when the core of your system is non-deterministic, its tools live on remote servers, and the output is a stream of natural language? Our team runs a multi-agent customer support system in production, and this talk is about how we keep it from breaking. Most of the heavy lifting is still done by the boring classics: unit and integration tests for all the code around the model, because that code is where most bugs actually live. The interesting part sits on top. We built an eval layer with a small golden dataset curated from real conversations, deterministic checks that verify what the agent decided to do (which sub-agent it routed to, which tool it called, whether the answer contains what it must), and LLM judges for the softer qualities that have no objective test. I'll show you what worked, what we had to throw away, and which kind of test catches which kind of failure, with regression stories from our CI pipeline.
MORE VOICES
Other speakers at code.talks
See them live — and 100+ more.
One ticket, every track. Nov 4–5, 2026 at Kinopolis Hamburg.
Get your ticket





