Build with AI agents and stay in control
I'm a three-time CTO and data executive, helping engineering teams build software with AI agents without giving up quality, control or measurement.
The method is the engineered AI-DLC. I built a working proof of concept with enterprise controls from the first story. The figures on this page come from its logs.
- ~30 minmedian from approved story to merged code
- ~7 minmedian agent pipeline time per story
- 17 of 19dispatches passed every gate first time
Figures describe the proof of concept. Method, scope and the full ledger are under Proof.
Move fast without sacrificing quality or value
Teams that get real value from AI agents don't have better models. They have better discipline: the right tool for each job, controls at every step and measurement from day one. These six principles build that discipline into delivery.
Problems are caught before code exists
Nothing is trusted on an agent's word
There is no pilot‑to‑production rewrite
The unit economics are measured
It scales on the control plane, not on headcount
Reliability starts with the right tool for each job: code, generative AI or machine learning
Many builders use LLM prompting for data-heavy work. This raises token costs and gives answers that can change from run to run, which a rules-based or statistical engine would not.
Code where the answer can be computed
When a question has one right answer that a rule can compute, code gives it every time, at no token cost and with a test to prove it. Calculations, matching and status logic belong here, not in prompts.
GenAI where interpretation is needed
Language models earn their cost on judgement over unstructured or ambiguous material, such as summarizing, classifying and reasoning over evidence. Every output cites its evidence so a person can check it.
ML when the data justifies it
Traditional, well-tested ML models are the right tool for business insight wherever the data exists, such as forecasting, segmentation and propensity. They are proven, explainable and cheap to run, and they are reached for before a language model.
Every change runs through four phases, with controls built into each one
People make the decisions and own the merge. Agents run the steps in between, inside a governance control plane that applies the controls in every phase.
Intake
A request becomes a story an agent can build without guessing.
- Who
- Product and technical leads approve requirements and stories before dispatch.
- Standard
- No dispatch until the coverage gate passes and a pre-dispatch review is done. All stories follow documented, pre-approved development standards.
Build
Agents construct in small batches, with no memory except what is written down.
- Who
- A fresh coder agent per story, in its own git worktree, orchestrated by a lead agent.
- Standard
- One story per dispatch. Handoffs go through files, never chat. Deny rules block destructive commands, and real secrets never enter a worktree.
Test and verify
Automated gates do the routine checking so human review can focus on judgement.
- Who
- Review agents run unit and regression tests, architecture review, QA against acceptance criteria and smoke tests.
- Standard
- Gates are non-optional. A failed gate returns the work to the agent automatically. Reviewers break the code on purpose before trusting a test.
Merge and operate
A person merges every change, and production problems return through the same pipeline.
- Who
- A human engineer reviews the test output and merges all new code that is being deployed to production. No agent can ever merge.
- Standard
- Branching rulesets are in place and enforced, and bugs are ranked and prioritized for fixes.
Decisions run through every phase, at three points
Nothing enters the queue before validation and review, so problems are caught before they get out of the gate.
When a choice is not covered by the story or existing decisions, the pipeline stops and asks the decision owner.
Every decision gets an ID as the work progresses, so the next agent and the next reviewer build on it.
Agents act according to instructions, and people decide what matters
Speed never comes at the cost of accountability, because every decision has a tier, an owner and a record.
Proceed
Do the work, and log it.
Choices inside the story that follow documented rules, including refactoring adjacent code so the story lands cleanly.
Naming conventions, extra tests
Recommend
Open a request for human review.
Present validated results for review and deployment. A person reviews every recommendation before anything ships.
New feature implementation
Escalate
Stop and ask the decision owner.
Changes that would break an existing decision, or ambiguous ones where no rule exists.
Changes to product scope, design, data definitions, access controls and security
Thirty minutes from approved story to merged code, and most of it is human review by design
Time from dispatch to merged on main, by story
Fourteen of fifteen gated stories, in minutes. One story that paused for a decision has no recorded split and is not shown; four dispatches are excluded from timing under the 30 September method. Agent time includes any automatic retry.
Delivery metrics
| Metric | Result |
|---|---|
| Lead time for changes | ~30 min median (traditional estimate: 1–2 days) |
| Deployment frequency | ~2.7 merges per active day; peak of 5 |
| Change failure rate | ~5%: one escaped severity-2 defect in 19 story merges |
| Time to restore | 9 min 40 s of pipeline time (the merge waited until morning) |
| Verification tax | 72% of elapsed time sits at the human review gate, including waiting |
| Control plane | Unchanged since 30 September while the product code grew 42%; now 35% of all code |
| Interventions | First 5 dispatches: 2 decision stops, 4 infrastructure stops. Next 14: 1 infrastructure stop. 2 automatic retries in total |
Pipeline live 24 September to 8 October 2026, with merges on seven days. Figures come from pipeline and git logs, never self-reported.
The approach is aligned with current industry best practice
Google Cloud's DORA team published The ROI of AI-assisted Software Development in 2026. Its findings line up with the build: AI amplifies what is already there, verifying AI output is a real cost, gains on simple new work far exceed gains on legacy code, productivity dips before it recovers, and freed capacity is better redeployed than cut. InfoQ's coverage summarizes the report.
The research prescribes non-optional automated gates, pre-commit hooks, heavy test automation, architecture decision records and small batches. The proof of concept has all five, and its ledger measures the verification tax directly.
A phased implementation program takes teams from baseline to scale
The engineered AI-DLC comes with a four-stage implementation program, and the first stage measures the team's baseline before anything changes. Durations are a proposal to test with the team, not a commitment.
- Days 1–30
Foundations
Measure the current DORA baseline. Stand up the control plane on the team's infrastructure. Adopt the story template and decision tiers with one pilot team.
- Days 31–60
Integrations
Connect intake to the team's real tools: tickets, chat and repositories. Run the first stories through the full pipeline on production code.
- Days 61–90
Agentic workflows
Widen to more teams. Automated gates become the default path, not the exception.
- Day 90 on
Governance and scale
Controls become automated gates where possible, metrics are reported upward and decision rights are settled.
What changes, and what to expect
- A dip comes first.DORA calls it the J-curve: productivity drops while the platform and habits are built. The proof of concept had a small one. The leader's job is to set expectations upward and plan for it.
- Spec quality becomes the job.When construction takes minutes, an unclear story is the bottleneck.
- Roles shift; they don't disappear.More time on design, review and decisions, less on typing and waiting for handoffs.
- Not every change suits the pipeline.Small, well-defined stories do. Sprawling legacy changes need breaking down first.
Twenty-five years in data science and engineering, now focused on generative AI at scale
I'm a technology and data executive with a PhD in astrophysics. I have been CTO three times, across multiple industries, and I led digital transformation, cloud transformation and the big-data wave hands-on.
I am convinced generative AI is a bigger shift than any of those. I built the proof of concept myself to test that conviction against real engineering standards, and I publish the numbers because I was trained to work from evidence.
I'm a Toronto native with experience in Canadian, American and European markets and industries.
- PhD, astrophysicsEvidence first: figures carry their source and assumptions.
- About 25 years in data science and engineeringFrom hands-on data work to executive leadership.
- Three-time CTOAcross multiple industries.
- Based in Oakville, OntarioWorking across the Greater Toronto Area and remotely.
It’s time to amplify your team’s impact
I'm open to executive and advisory roles. If your engineering organization wants AI agents in its delivery process with the controls and the measurement to match, I would like to hear about it.
| CTO | Full-time leadership of engineering, data and AI delivery. |
| Interim or transformation CTO | Lead a defined change, such as moving a team to agent-based delivery. |
| Fractional CTO | Part-time technology leadership for a company that needs the judgement without the full role. |
| Chief Data Officer or Chief AI Officer | Own data strategy, AI governance and measurement. |
Email: [email protected]
LinkedIn: linkedin.com/in/moranjane
A walkthrough of the proof of concept is available on request.