Clawd, the reliability plumber

Flattening the Agent
Reliability Curve

How coding agents are re-writing reliability

Ezo Saleh
Co-Founder / CTO, orra

Flattening the curve

The Reliability Problem

Complexity → Reliability → The Cliff
Traditional automation: Works great for known patterns, falls apart at the edge of capability

Flattening the curve

Orchestration v1:
Human as Exception Handler

The Traditional Approach

  • Maximum automation
  • Minimal human friction
  • Escalate when uncertain
  • Intervene on failure
  • Define workflows statically
↓
Compounds mistakes before you can intervene

Flattening the curve

The Inversion

Human oversight is the
primary execution model,
not exception handling

Old: Autonomy Despite Humans

Human intervention = failure

New: Autonomy Within Humans

Human checkpoints = decision points

Flattening the curve

Flattening the Reliability Curve

Complexity → Reliability → - - - Traditional ━━━ Human-Guided
Human oversight extends the reliability frontier into territory where pure automation breaks down

Flattening the curve

Beyond Coding →
Personal Delegation

Claude Code's Evolution

Started: Coding assistant

Became: Multi-agent delegation platform

Enables: Personal orchestration for any task

Write code • Analyze systems • Create content •
Manage tasks • Research topics • Deploy infrastructure

How It Works

The Architecture That Makes This Possible

How it works

The Dual-Ended System

LLM
⟷
Client

↓

Both tracking state simultaneously

LLM Side:
  • Context management
  • Reasoning & planning
  • Tool selection
Client Side:
  • Memory persistence
  • Checkpoint handling
  • Permission control

How it works

Memory: The 200K Token Budget

200K Tokens per agent
~800K Characters
∞ With sub-agents
What counts: User messages + Tool calls + Tool results + History + System prompts

How it works

Sub-Agents = Infinite Context

// Main agent running low on context Main Agent [190K tokens used] ↓ Task("Explore authentication patterns") ↓ Sub-Agent [Fresh 200K budget] → Search codebase → Read auth files → Analyze patterns → Return concise summary (2K tokens) ↓ Main Agent [190K → 192K] ← Receives summary, not full context!
Offload expensive work,
receive compact results

How it works

Context Engineering Strategies

Proactive (Agent Chooses)

  • Grep before Read
  • Read with offset/limit
  • Spawn Task agents
  • Strategic tool selection

Automatic (SDK Level)

  • Context compaction
  • Summarization
  • Priority-based truncation

Agents as
Markdown Files

Agents as markdown files

Skills = Markdown + Front Matter

--- name: code-reviewer description: Review code for bugs and improvements --- # Code Review Guidelines Analyze the code for: - Security vulnerabilities - Performance issues - Code smells and anti-patterns - Test coverage gaps Provide actionable feedback with examples.
That's it. No Python classes.
No configuration files.
Just prompts.

Agents as markdown files

The Planner/Scaffold

High-Level Agent
Planner
↓ manages ↓
Operations
Extendability
Checkpoints
Skills
A markdown-based scaffold managing the entire agent lifecycle

Agents as markdown files

Checkpoints:
Human Comprehension

0
1
2
3
4
5
Pause. Review. Resume.
Traditional:
Fault tolerance
(retry failed steps)
Claude Code:
Human comprehension
(verify then continue)

Agents as markdown files

Permission Modes:
Calibrated Control

ask

Request approval
before execution

auto

Execute tools
automatically

never

Block specific
operations

Fine-grained control: per-tool, per-category, with custom hooks

Making Skills
Deterministic

Making skills deterministic

Almost Deterministic Skills

What Makes Skills Predictable?

  • Clear markdown instructions
  • Structured input/output expectations
  • Constraint-based reasoning
  • Validation through human checkpoints

Not "truly deterministic" like traditional code

But predictable enough for production

Making skills deterministic

Both Ends Tracking
Simultaneously

LLM Context window Reasoning state Tool results Planning Client Conversation log Checkpoints Permissions Memory
State synchronization between LLM reasoning and client memory creates robust reliability

Why This Matters

Re-Writing Reliability

Why this matters

Reliability at the
Edge of Capability

Traditional Automation ✗

Known patterns ✓ | Unknown patterns ✗

Human-Guided Agents ✓

Known patterns ✓ | Unknown patterns ✓ | Edge cases ✓

Why this matters

Accessibility:
Markdown > Complex Frameworks

Traditional Agent Frameworks

  • Python classes
  • Configuration files
  • Deployment pipelines
  • Testing infrastructure
  • Monitoring setup

Claude Code

  • Markdown file
  • Front matter
  • Natural language
  • Done ✓
Creating agents shouldn't require
engineering infrastructure

Why this matters

Real Example:
Complex Refactoring

Task: "Refactor authentication to use OAuth2 across 8 services" // Agent reasoning with checkpoints: 1. Explore current auth implementation → Checkpoint: "Found 3 different auth patterns, proceed?" 2. Plan migration strategy → Checkpoint: "Here's the approach—make sense?" 3. Implement OAuth2 client library 4. Migrate services one by one → Checkpoint after each: "Service N migrated, tests pass?" 5. Update docs and deploy
Real-time reasoning. Human validates at decision points. No predefined workflow.

Why this matters

Where Traditional
Automation Fails

  • Novel situations that don't fit templates
  • Ambiguous requirements needing clarification
  • Complex reasoning across multiple contexts
  • Edge cases not in training data
  • Tasks requiring creative problem-solving
  • Situations needing ethical judgment
= Where Claude Code thrives

Why this matters

Where Claude Code Thrives

  • Complex one-off tasks
  • Exploratory work (research, analysis)
  • Reasoning-heavy problems
  • Tasks with multiple valid approaches
  • Learning new domains
  • High-stakes decisions needing oversight
Key insight: Not replacing automation—complementing it with supervised intelligence

Why this matters

Production Pattern:
Start Conservative, Scale Trust

Phase 1: Learning
Use ask mode • Understand what the agent does
Phase 2: Trusting
Move to auto • Proven patterns run faster
Phase 3: Protecting
Use never + hooks • Validation gates for critical ops

Why this matters

What This Enables

Boring > Flashy
Reliability through oversight isn't sexy. But it works.
Delegation Becomes Natural
Like working with a fast junior developer who needs guidance.
Markdown as Interface
Creating agents shouldn't require infrastructure.
Human + AI > Either Alone
Extending what humans can accomplish through supervised reasoning.

Why this matters

Key Takeaways

1. Invert the Model
Human oversight as primary, not exception handling
2. Flatten the Curve
Extend reliability into territory where automation breaks
3. Dual-Ended Architecture
LLM + Client both tracking state simultaneously
4. Markdown as Agents
Skills, scaffolds, and prompts—no infrastructure needed
5. Almost Deterministic
Predictable enough for production through constraints + validation
6. Re-Writing Reliability
Making agentic workflows accessible like nothing before

Questions?

Claude Code docs:
code.claude.com/docs

Thank you!