Engineering

May 8, 2026

9 min

Image

Building reliable agents: lessons from 2 billion task executions

After processing over 2 billion agent tasks, we've learned what makes agents fail, what makes them succeed, and the architectural patterns that separate toy demos from production systems.

TL;DR

  • Context management is the #1 driver of agent reliability (38% of failures)

  • Structured ReAct loops improve success rates by 37 percentage points

  • Every tool call must have timeout + retry + fallback — no exceptions

  • Memory retrieval precision beats breadth every time

  • Human approval gates should be granular, not binary on/off

When we started building Forge, we thought reliability would be about choosing the right model. We were wrong. After processing 2.1 billion agent tasks across thousands of production workspaces, we've learned that model quality is just one variable in a complex system where architecture, prompting strategy, error handling, and memory management all play equally critical roles.


In this post, I'll share the five most impactful lessons we've learned — with data from our production systems and concrete implementation patterns you can apply today.


Context overflow kills agents silently


The number one cause of agent failure in our data is context overflow — 38.2% of all failures. What makes this particularly insidious is that it often fails silently. The model doesn't throw an error; it just starts hallucinating, repeating itself, or forgetting earlier parts of its task.
The fix requires proactive context management: summarizing earlier turns, using structured memory instead of raw conversation history, and designing tasks to be completable within a bounded context window.


The fix requires proactive context management: summarizing earlier turns, using structured memory instead of raw conversation history, and designing tasks to be completable within a bounded context window.


"Context overflow doesn't announce itself. Agents just start confusing themselves until the task quietly fails."
— Jordan Wu, Lead AI Researcher


Structured reasoning dramatically improves success rates


Agents prompted to explicitly think → plan → act → observe before taking action achieve 40% higher task completion rates than agents with simple "do this" instructions. The improvement is consistent across task types, model families, and complexity levels.

Without structured reasoning

0.4%

With Forge ReAct loop

0.4%

Without structured reasoning

0.4%

With Forge ReAct loop

0.4%

Implementation: the ReAct loop in Forge

Here's how we implement structured reasoning in the Forge SDK. The key insight: make reasoning steps explicit and observable — not just a hidden system prompt.

Image
Image
Image
Image
Image
Image

Tool timeout handling is non-negotiable


24.7% of failures in our data are caused by tool timeouts. Agents without graceful timeout handling tend to either hang indefinitely or abort the entire task when a single tool call fails.
Every tool should have a hard timeout, a soft retry budget, and a degraded fallback. When a web search fails, the agent should proceed with reduced confidence rather than aborting.


Memory quality beats quantity


More memory doesn't mean better agents. In our experiments, agents with curated semantic memory (top-5) outperformed agents with exhaustive recall on complex tasks. Flooding context with tangential memories hurts.

Stop typing the boring parts. Start shipping the good parts.

Join 12,000+ teams using Forge to ship faster, cleaner, and with fewer 3 AM pages. Free for solo devs forever.

Image
Image
Image

Stop typing the boring parts. Start shipping the good parts.

Join 12,000+ teams using Forge to ship faster, cleaner, and with fewer 3 AM pages. Free for solo devs forever.

Image
Image
Image

Stop typing the boring parts. Start shipping the good parts.

Join 12,000+ teams using Forge to ship faster, cleaner, and with fewer 3 AM pages. Free for solo devs forever.

Image
Image
Image
Logo

The AI engineering platform. Made with too much coffee in Brooklyn, Berlin, and Bengaluru.

Media
Media
Media
Media
Media
Media
Media
Media
Media
Media

© 2026 Forge Labs, Inc.

Image
Logo

The AI engineering platform. Made with too much coffee in Brooklyn, Berlin, and Bengaluru.

Media
Media
Media
Media
Media
Media
Media
Media
Media
Media

© 2026 Forge Labs, Inc.

Image
Logo

The AI engineering platform. Made with too much coffee in Brooklyn, Berlin, and Bengaluru.

Media
Media
Media
Media
Media
Media
Media
Media
Media
Media

© 2026 Forge Labs, Inc.

Image

Create a free website with Framer, the website builder loved by startups, designers and agencies.