Case Study - From manual reconciliation to agents you can check

A data platform, agents that build its pipelines, and a data agent that answers from it. At every layer, whoever does the work does not check it.

Client
A regulated insurance business
Year
Service
Agentic data platform
Data flows from internal systems into pipelines built by agents, through a daily data platform, to a data agent, and on to actuaries. Their real questions become evals for the agents.
Internal systems and cedant data

The inputs to every regulatory filing.

Agents build the pipelinesIngestion skillsRouting skills

A separate agent writes the checks from the spec.

The data platform runs them daily

Every run is versioned, reproducible and auditable.

A data agent answers from the dataGuard skills

It checks the data first, and stops if it looks wrong.

Actuaries review in at most an hour

Down from 2-3 weeks of manual reconciliation.

Shared by every agent

Written once with your team, and reused across agents.

Domain skillsAudit-check skillsEvals on real questions

The problem

Cedant files arrive several times a day. Before each regulatory filing, actuaries spent 2-3 weeks reconciling them against internal systems by hand. Every cycle meant re-work, and none of it was reproducible.

The platform

We built a data platform on Databricks with a squad of three. It runs the ingestion jobs and the checks for them.

2-3 weeks of manual reconciliation became a daily run, needing at most an hour of actuarial review. Every run is versioned, reproducible and auditable.

Agents that build the pipelines

Agents now write the ingestion pipelines. The risk is a builder that grades its own homework.

So a separate agent writes the checks from the spec. Every pipeline is tested end to end on real jobs before it ships. A pipeline that took 3 days now takes 30 minutes.

A data agent that answers from the data

A data agent answers questions on the curated data. Early on, it was confident and wrong. It said a cost was rising when it had fallen every year.

Guard skills now check the shape of the data first. The agent stops when something does not look right.

Written rules did not fix everything. Three rounds of prose rules never reached the SQL the agent generated. One example query did. We read the generated SQL, not just the answer.

Skills written once, reused by every agent

Every agent here follows skills: short written procedures for one job. We write and review them with your team. Evals on real questions check they keep working.

One skill can serve many agents and teams. The audit checks are skills too, so every agent runs the same ones.

  • Ingestion. Skills take a ticket from spec to code ready for review. A check runs after every change.
  • Routing. Skills decide where each incoming file goes, and test that decision.
  • Domain. Skills teach each agent the reinsurance terms, and turn a business ask into a spec.
  • Data agent. A skill checks the data's shape first, then answers only from queries we have checked.

What we did

  • Databricks
  • Domain mapping
  • Agentic pipeline generation
  • Evals and benchmarks
  • Guard skills
  • Token metering
To build a checked pipeline
3 days → 30 min
Filing reconciliation
2-3 weeks → daily
Versioned and auditable
Every run

An agent repeats your data's mistakes, at machine speed.

Book a call. We'll find where that happens for you.