Developersummit
  • HOME
  • CALL FOR PAPERS
  • BUY TICKETS
  • CONTACT
  • INSIGHTS
  • ONDEMAND
saltmarch

GIDS news media, articles, insights and virtual events educate and illuminate its audiences so they can be fully prepared to deal with the new realities at work and in their professions.

Saltmarch On-Demand
Media

Our Experts

Videos On Demand

Insights

Call for Papers

Connect

About Us

Privacy Policy

Terms & Conditions

Code of Conduct

Contact Us

Subscribe to Developersummit

Get the latest event updates, and insights from today's leading voices.

© 2026-2027 Saltmarch. All rights reserved.

What Breaks When Research Agents Go to Production?
RegisterTwitterLinkedInFacebook

< session />

What Breaks When Research Agents Go to Production?

Thu, January 1AI-NativeArchitecture & Distributed Systems

The hardest part of building production research agents is not generating impressive answers. It is controlling the environment in which they reason.

Many AI prototypes perform well on carefully selected examples: retrieve context, call a model, and generate a response. Once these systems enter real expert workflows, the failure modes become more subtle. Systems may retrieve relevant but incomplete evidence, include too much context and bury important signals, route questions to the wrong source, pass weak evidence into later stages, lose constraints across long workflows, or produce confident conclusions that are difficult to verify.

This session examines what breaks when research and reasoning workflows move from demonstrations into production. Drawing on experience building AI systems for knowledge-intensive and regulated environments, it focuses on the engineering discipline required around the model, including context engineering, tool boundaries, orchestration, evaluation loops, observability, citation-backed synthesis, human review, latency, cost, and traceability.

A central theme of the talk is that larger context windows do not eliminate the need for context discipline. Production systems must determine what context each step should see, what should be excluded, how evidence should be ranked, when missing context should trigger reflection, and how intermediate outputs should be carried forward without affecting subsequent steps.

Attendees will learn a practical framework for identifying and strengthening common failure modes in production research agents, including wrong context, missing context, excessive context, weak grounding, poor tool selection, weak evidence handoff, insufficient review boundaries, and outputs that cannot be traced back to source material. The session focuses on designing the controls, context flows, evaluation loops, and recovery mechanisms that make research workflows suitable for expert use.

What You Will Learn:

  • Common failure modes that emerge when research agents move from prototypes into production workflows
  • How context engineering, tool boundaries, evidence handling, and review processes influence system behaviour
  • A practical framework for identifying and strengthening weaknesses in production research and reasoning workflows

Who Should Attend:

  • AI engineers
  • Applied AI practitioners
  • Platform engineers
  • Staff and principal engineers
  • Architects building agent-based systems
  • Teams developing research or knowledge-intensive workflows
  • Technical leads responsible for AI system quality and trustworthiness

< speaker_info />

About the speaker

Sarang Kulkarni

Sarang Kulkarni

Principal Consultant, Thoughtworks

Sarang Kulkarni is a Principal Consultant at Thoughtworks with over 14 years of experience as a polyglot developer, spanning software development, data engineering, DevOps, and AI. He currently leads a healthcare client account, spearheading Generative AI initiatives that accelerate drug discovery processes by making decades of study data more accessible and actionable.

Recently, Sarang joined Thoughtworks' Global AI Service Development team, contributing to the organization's evolving AI strategy and connecting client delivery with broader strategic initiatives.

A continuous learner at heart, Sarang consistently explores emerging technologies and approaches. Beyond client work, he is an O'Reilly Media trainer specializing in productionizing RAG applications and has spoken at numerous conferences, where he is recognized for blending technical depth with practical, real-world guidance.

Related Talks

When Healthy Agents Fail: 12 Patterns from Production

When Healthy Agents Fail: 12 Patterns from Production

Tuhin Sharma, Soham Dutta
Scaling Developer Infrastructure for an Agentic World

Scaling Developer Infrastructure for an Agentic World

Daniel Nadasi
Designing Codebases for Coding Agents

Designing Codebases for Coding Agents

Ragunath Jawahar

On-Demand Talks

All On-Demand »