Written, not implied
A spec an agent can run is also a spec a regulator can read. The two requirements turn out to be the same requirement.
AI agents do a great deal of the building now. That changes how fast the work moves — it does not change who is accountable for what ships.
Every phase is accelerated by agents and closed by a named human. Nothing advances because it looks finished; it advances because someone signed for it.
| Phase | What happens | Human gate |
|---|---|---|
| 01 · Discover | Use-case triage against value and feasibility; data audit; agents sweep your documents and systems to map what already exists. | BUSINESS CASE SIGNED |
| 02 · Design | Solution architecture, evaluation design, and guardrail specification — written so both agents and auditors can execute it. | ARCHITECTURE APPROVED |
| 03 · Build | Agent teams implement, test, and document in parallel; engineers direct the work and review every change. | EVAL SUITE PASSING |
| 04 · Evaluate | Offline evals, red-teaming, and user acceptance with your subject-matter experts on real data. | QUALITY BAR MET |
| 05 · Deploy | Staged rollout with human-in-the-loop first; integration hardening; operations training. | READINESS REVIEW |
| 06 · Operate | Monitoring, drift detection, retraining, and cost management, run as a service. | MONTHLY VALUE REVIEW |
Five rules that make agent-accelerated delivery safe enough to put in front of an auditor — and fast enough to be worth doing.
A spec an agent can run is also a spec a regulator can read. The two requirements turn out to be the same requirement.
“It seems to work” is not a status. The eval suite is the status, and it runs on every change.
Retrieval and citation are structural. A confident answer with no source is a defect, not a feature.
Speed comes from the agents. Accountability comes from a named person at every phase boundary.
Every phase ends with something running and a number attached to it. We transfer knowledge as we go: your team owns the outcome, not a dependency on us.
| Phase | What we do together | You get | Typical timeline |
|---|---|---|---|
| Sprint zero | Use-case prioritisation workshop, data readiness assessment, and a value model for each candidate. | A ranked portfolio and a pilot plan | 2 WEEKS |
| Pilot | One use case built on your data against agreed KPIs — agents accelerate, your team stays in the loop throughout. | A working system and a measured result | 4–8 WEEKS |
| Production | Hardening, integration into your systems and workflows, security review, and operations training. | A deployed system your team runs | 8–12 WEEKS |
| Operate & scale | MLOps as a service — monitoring, retraining, cost control — while the next use cases enter the pipeline. | A compounding AI portfolio | ONGOING |
Watch The pilot’s KPI is agreed before a line of code is written. If we can’t name the number, we haven’t found the use case yet.
The engineers who design the model also wire the devices, build the pipeline, and run it in production. No handoff, no seam for the work to fall through.
We build internal capability as we go. The measure of a good engagement is that your team can run and extend what we built without calling us.
A baseline before, a measurement after, and a monthly value review once it is running. Everything else is decoration.
Sprint zero is small, fixed, and ends with a plan you could hand to another firm if you wanted to. That is rather the point.