Flattening the Agent Reliability Curve
How coding agents are re-writing reliability
Ezo Saleh
Co-Founder / CTO, orra
We're exploring a fundamentally different approach to agent reliability. This talk is about what, how, and why coding agents are re-writing what's possible.
Flattening the curve
The Reliability Problem
Complexity →
Reliability →
The Cliff
Traditional automation: Works great for known patterns, falls apart at the edge of capability
Traditional agent systems hit a reliability cliff. They work beautifully for known patterns but fall apart when dealing with novel situations, edge cases, or complex reasoning. The more complex the task, the less reliable they become.
Flattening the curve
Orchestration v1: Human as Exception Handler
The Traditional Approach
Maximum automation
Minimal human friction
Escalate when uncertain
Intervene on failure
Define workflows statically
↓ Compounds mistakes before you can intervene
Previous frameworks like Airflow or Prefect treated humans as exception handlers. The goal was maximum automation with minimal human friction. But when agents are doing complex reasoning, mistakes compound before you can intervene.
Flattening the curve
The Inversion
Human oversight is theprimary execution model , not exception handling
Old: Autonomy Despite Humans
Human intervention = failure
New: Autonomy Within Humans
Human checkpoints = decision points
Claude Code inverts this completely. Human oversight isn't a constraint—it's central to the architecture. You get autonomy within human-guided boundaries, not despite them.
Flattening the curve
Flattening the Reliability Curve
Complexity →
Reliability →
- - - Traditional
━━━ Human-Guided
Human oversight extends the reliability frontier into territory where pure automation breaks down
This is the core value proposition: flattening the reliability curve. With human oversight woven in, agents remain reliable even at the edge of capability—handling novel situations, edge cases, and complex reasoning that would break traditional automation.
Flattening the curve
Beyond Coding → Personal Delegation
Claude Code's Evolution
Started: Coding assistant
Became: Multi-agent delegation platform
Enables: Personal orchestration for any task
Write code • Analyze systems • Create content • Manage tasks • Research topics • Deploy infrastructure
Claude Code started as a coding assistant but has evolved into a personal multi-agent delegation platform. It's not just about code anymore—it's about orchestrating complex tasks across any domain with human supervision.
How It Works
The Architecture That Makes This Possible
Let's dive into the technical architecture. How does Claude Code actually achieve this reliability?
How it works
The Dual-Ended System
↓
Both tracking state simultaneously
LLM Side:
Context management
Reasoning & planning
Tool selection
Client Side:
Memory persistence
Checkpoint handling
Permission control
Here's what makes Claude Code different: it's a dual-ended system. The LLM handles reasoning and context, while the client handles memory persistence, checkpoints, and permissions. Both ends work together, tracking state simultaneously.
How it works
Memory: The 200K Token Budget
200K
Tokens per agent
~800K
Characters
∞
With sub-agents
What counts: User messages + Tool calls + Tool results + History + System prompts
Every agent starts with a 200K token budget—that's roughly 800K characters. This includes everything: messages, tool calls, results, history. It's a massive context window that enables long-running sessions.
How it works
Sub-Agents = Infinite Context
Main Agent [190K tokens used]
↓
Task("Explore authentication patterns")
↓
Sub-Agent [Fresh 200K budget]
→ Search codebase
→ Read auth files
→ Analyze patterns
→ Return concise summary (2K tokens)
↓
Main Agent [190K → 192K]
← Receives summary, not full context!
Offload expensive work, receive compact results
When the main agent runs low on context, it spawns sub-agents with fresh 200K budgets. The sub-agent does expensive operations and returns only a concise summary. This enables unlimited exploration without exhausting context.
How it works
Context Engineering Strategies
Proactive (Agent Chooses)
Grep before Read
Read with offset/limit
Spawn Task agents
Strategic tool selection
Automatic (SDK Level)
Context compaction
Summarization
Priority-based truncation
Context engineering happens at two levels. Proactively, the agent uses strategies like grepping before reading and spawning sub-agents. Automatically, the SDK compacts old conversation turns and manages truncation.
Agents as Markdown Files
Now the most elegant part: how agents are defined as simple markdown files with front matter.
Agents as markdown files
Skills = Markdown + Front Matter
name: code-reviewer
description: Review code for bugs and improvements
Analyze the code for:
- Security vulnerabilities
- Performance issues
- Code smells and anti-patterns
- Test coverage gaps
Provide actionable feedback with examples.
That's it. No Python classes. No configuration files. Just prompts.
Skills are custom agents defined entirely in markdown with YAML front matter. A markdown file with instructions. No code. No deployment. Just prompts that guide the agent's reasoning. Creating agents is as easy as writing documentation.
Agents as markdown files
The Planner/Scaffold
High-Level Agent Planner
↓ manages ↓
Operations
Extendability
Checkpoints
Skills
A markdown-based scaffold managing the entire agent lifecycle
There's a high-level agent planner—a scaffold defined in markdown—that manages operations, extendability, checkpoints, and skills execution. It orchestrates the entire agent lifecycle through markdown instructions.
Agents as markdown files
Checkpoints: Human Comprehension
Pause. Review. Resume.
Traditional: Fault tolerance(retry failed steps)
Claude Code: Human comprehension(verify then continue)
Checkpoints are about human comprehension, not fault tolerance. The agent does work, you pause, review what it accomplished, verify it's correct, then resume. It transforms long tasks from opaque automation into supervised reasoning sessions.
Agents as markdown files
Permission Modes: Calibrated Control
ask
Request approval before execution
auto
Execute tools automatically
never
Block specific operations
Fine-grained control: per-tool, per-category, with custom hooks
Permission modes give fine-grained control over autonomy. Ask for high-stakes tasks. Auto for confident workflows. Never to block destructive operations. Plus custom hooks for validation gates. This is calibrated control, not all-or-nothing.
Making Skills Deterministic
How do we make agent behavior predictable and reliable?
Making skills deterministic
Almost Deterministic Skills
What Makes Skills Predictable?
Clear markdown instructions
Structured input/output expectations
Constraint-based reasoning
Validation through human checkpoints
Not "truly deterministic" like traditional code
But predictable enough for production
Skills aren't truly deterministic like traditional code, but they're predictable enough for production. Clear markdown instructions, structured expectations, constraints, and human validation create reliable behavior patterns.
Making skills deterministic
Both Ends Tracking Simultaneously
LLM
Context window
Reasoning state
Tool results
Planning
Client
Conversation log
Checkpoints
Permissions
Memory
State synchronization between LLM reasoning and client memory creates robust reliability
The LLM tracks reasoning state and context. The client tracks conversation log, checkpoints, and permissions. They synchronize state continuously, creating robust reliability through redundancy.
Why This Matters
Re-Writing Reliability
Let's talk about why this architecture matters and what it enables.
Why this matters
Reliability at the Edge of Capability
Traditional Automation ✗
Known patterns ✓ | Unknown patterns ✗
Human-Guided Agents ✓
Known patterns ✓ | Unknown patterns ✓ | Edge cases ✓
Traditional automation handles known patterns but fails at edges. Human-guided agents extend reliability into novel situations, edge cases, and complex reasoning that would break pure automation.
Why this matters
Accessibility: Markdown > Complex Frameworks
Traditional Agent Frameworks
Python classes
Configuration files
Deployment pipelines
Testing infrastructure
Monitoring setup
Claude Code
Markdown file
Front matter
Natural language
Done ✓
Creating agents shouldn't require engineering infrastructure
Traditional frameworks require Python classes, configs, deployment, testing, monitoring. Claude Code requires a markdown file with front matter. This makes agent creation accessible to anyone who can write instructions.
Why this matters
Real Example: Complex Refactoring
Task: "Refactor authentication to use OAuth2 across 8 services"
1. Explore current auth implementation
→ Checkpoint: "Found 3 different auth patterns, proceed?"
2. Plan migration strategy
→ Checkpoint: "Here's the approach—make sense?"
3. Implement OAuth2 client library
4. Migrate services one by one
→ Checkpoint after each: "Service N migrated, tests pass?"
5. Update docs and deploy
Real-time reasoning. Human validates at decision points. No predefined workflow.
Here's a real example: refactoring auth across multiple services. The agent explores code, plans strategy, implements changes, migrates services incrementally. At each decision point, there's a human checkpoint. No predefined workflow—just supervised reasoning through a complex task.
Why this matters
Where Traditional Automation Fails
Novel situations that don't fit templates
Ambiguous requirements needing clarification
Complex reasoning across multiple contexts
Edge cases not in training data
Tasks requiring creative problem-solving
Situations needing ethical judgment
= Where Claude Code thrives
Traditional automation maxes out on novel situations, ambiguous requirements, complex reasoning, edge cases, creative problems, and ethical judgment. These are exactly where Claude Code with human guidance thrives.
Why this matters
Where Claude Code Thrives
Complex one-off tasks
Exploratory work (research, analysis)
Reasoning-heavy problems
Tasks with multiple valid approaches
Learning new domains
High-stakes decisions needing oversight
Key insight: Not replacing automation—complementing it with supervised intelligence
Claude Code excels at complex one-off tasks, exploratory work, reasoning-heavy problems, and high-stakes decisions. It's not replacing traditional automation—it's complementing it with supervised intelligence for tasks that require it.
Why this matters
Production Pattern: Start Conservative, Scale Trust
Phase 1: Learning
Use ask mode • Understand what the agent does
Phase 2: Trusting
Move to auto • Proven patterns run faster
Phase 3: Protecting
Use never + hooks • Validation gates for critical ops
Here's the production pattern we've learned: Start conservative with ask mode—understand what the agent does. Move to auto for proven patterns. Use never mode and hooks for critical operations. Scale trust gradually.
Why this matters
What This Enables
Boring > Flashy
Reliability through oversight isn't sexy. But it works.
Delegation Becomes Natural
Like working with a fast junior developer who needs guidance.
Markdown as Interface
Creating agents shouldn't require infrastructure.
Human + AI > Either Alone
Extending what humans can accomplish through supervised reasoning.
This architecture enables something important. Boring over flashy—reliability works. Delegation becomes natural—like working with a junior developer. Markdown as interface—no infrastructure needed. And fundamentally, human plus AI is greater than either alone.
Why this matters
Key Takeaways
1. Invert the Model
Human oversight as primary, not exception handling
2. Flatten the Curve
Extend reliability into territory where automation breaks
3. Dual-Ended Architecture
LLM + Client both tracking state simultaneously
4. Markdown as Agents
Skills, scaffolds, and prompts—no infrastructure needed
5. Almost Deterministic
Predictable enough for production through constraints + validation
6. Re-Writing Reliability
Making agentic workflows accessible like nothing before
Six takeaways: Invert the model. Flatten the curve. Dual-ended architecture. Markdown as agents. Almost deterministic through constraints. Re-writing reliability for accessible agentic workflows.
Questions?
Claude Code docs:code.claude.com/docs
Thank you! I'd love to hear your questions and discuss how you might apply these patterns in your own work.