New release : CTI Report - Pharmaceutical and drug manufacturing 

                 Download now

Industrializing AI-Assisted Development: What We Changed and Why

Over the past year the way we build software internally has changed shape. Code is largely produced by agents, and what a developer contributes has moved somewhere else. Along the way we built an internal framework, iagen-dev, to make that shift repeatable instead of dependent on who happens to be driving.

This article is the reasoning behind it. Not the installation procedure, not our internal tooling: the why. Most of it should transfer to any organization trying to do this seriously.


1. The posture shift

This is a change in the job itself, and it is not optional. Your trade is no longer writing code. It is:

  • Specifying: producing a good, stable, testable statement of the need.
  • Framing the agents: rules, context, the right task at the right size, in the right model.
  • Putting verification and guardrails in place: tests, linters, CI, reviews. You build the harness, the agents work inside it.
  • Running the project: sequencing, arbitrating, keeping scope under control.
  • Holding a product view of your own project: what it is for, who benefits, what you deliberately will not build, when you stop.
  • Owning the whole lifecycle: deployment, monitoring, sustainment, end of life.

Your unit of value is no longer source code. It is a good, stable, testable specification that a team of agents implements. You express the need, you review the spec, then you verify functionally, on end-to-end behavior and real runs, not line by line.

Nobody is demoted by this. The work moves from writing syntax to designing constraints, arbitrating trade-offs and auditing systems. That is a harder job than the one it replaces.

Two consequences people consistently underestimate. First, the output is ordinary software: compiled Go, a Vue/Nuxt frontend, built, tested, deployed on Kubernetes. No LLM at runtime, no magic. What changed is the production method, not the artifact. Second, the bottleneck moves to you. When agents struggle it is almost never executed, it is definition. “The problem is behind the keyboard” is the most frequent root cause we see.

BEFORE NOW Requirements, half-written Design, mostly in your head WRITING CODE where the hours went the work moves up SPECIFYING framing the agents building the harness verifying functionally owning the lifecycle designing constraints, arbitrating trade-offs Code, generated ordinary software: compiled, tested, deployed
Where the value moves. What you are judged on stops being source code, and becomes the specification and the harness around it.

2. The finding assumption: safety engineering

The framework borrows its structure from safety engineering, as practiced in avionics, automotive and medical software. Its founding assumption is one sentence:

The entity producing the code is not reliable. It makes mistakes. So you put a harness around it.

That was already true of humans. It is true of agents. Well harnessed, an unreliable producer yields a reliable result. This is not a new idea, it is why aviation software is trustworthy despite being written by people who have bad days.

Everything in iagen-dev is harness: linters, compiled languages, tests, CI, per-project rule files, multi-agent review. Nothing in it assumes the agent is good. It assumes the agent is productive and fallible, and builds the checks accordingly.

CROSS-MODEL REVIEW CI GATES: the last resort, not the discovery TESTS / TDD: you define the contract ZERO-DIFF LINTING TYPES & COMPILE THE PRODUCER: HUMAN OR AGENT assumed unreliable. It makes mistakes. Well-harnessed, an unreliable producer yields a reliable result.
The harness. Nothing in it assumes the producer is good. It assumes the producer is productive and fallible.

That framing settles a lot of arguments before they start. “Should we trust AI-generated code?” is the wrong question. You do not trust human-generated code either. You test it, lint it, review it, and run it in an environment that survives its failures. Apply the same discipline, calibrated to a producer that is faster, cheaper, tireless and differently wrong.

It also has a belief that comes back every time a language choice is discussed: that picking a memory-safe language removes the class of problem. It doesn't. Cloudflare's outage of 18 November 2025 is a good public example, and their postmortem is worth reading in full. A database permissions change made a query return duplicate columns, the Bot Management configuration file doubled in size, it crossed a hardcoded limit of 200 features against about 60 in use, and the proxy handling it, written in Rust, called unwrap() on an error and panicked. Memory safety is not correct.

The harder version of the same point is that a correct implementation of a wrong specification fails just as hard. In 1993 year A320 landed at Warsaw on one gear in a crosswind, on a wet runway. Ground spoilers required 6.3 tons on each main gear or wheels spinning above 72 knots, and aquaplaning meant neither condition was met. The braking logic remained inhibited for about nine seconds and the aircraft ran off the runway, killing two people. The software did exactly what its specification said. The specification's definition of being on the ground does not survive contact with that landing.

That is why the spec is the artifact that gets reviewed. A harness is only as good as the assumptions encoded in it, and no language checks those for you.

3. Why a framework at all

Developer adoption of AI tooling is near universal while trust is falling. In the Stack Overflow Developer Survey 2025, across 49,000 respondents, 84% use or plan to use AI tools, up from 76% the year before. In the same survey 46% do not trust the accuracy of AI output, up from 31% in the 2024 edition, 66% name “AI solutions that are almost right, but not quite” as their top frustration, and 72% say vibe coding is not part of their professional work.

84%use or plan to use AI tools, up from 76%
46%do not trust its accuracy, up from 31%
66%name «almost right, but not quite» as their top frustration
~⅓of improvement requests are about understanding the codebase

This disappointment is legitimate and well documented. It is also misattributed.

The top improvement developers ask for is better contextual understanding. In Qodo's State of AI Code Quality 2025, a third of all improvement requests are about the tool understanding the codebase, the team's norms and the project structure, and half the developers reporting a context gap work in organizations of ten people or fewer. This is not a big-company problem.

The concrete experience is familiar to anyone who has tried. On an existing service, the AI generates code that compiles but ignores the project's error patterns, reinvents existing abstractions, reaches for third-party libraries instead of the standard library, and writes tests that test the implementation rather than the behavior. It works but does not integrate, and you spend more time fixing it than you would have spent writing it. The rational conclusion is that AI does not work, when in fact it is the absence of a frame that does not work.

Everything the framework ships (per-project rule files, the documentation model, the lint gates, the project tiers) exists to supply that missing frame. It is distributed as a versioned skill pack that agents fetch and self-update from, so a fix found on one project propagates to the others.

4. Specifying is coding

The spec is no longer a preliminary document you write to satisfy a process. It is the artifact that determines the result.

This is a shared observation among the people doing it, not a surprise: a spec co-written with an agent, interactively, comes out better than what an experienced developer produces alone. The agent challenges assumptions, surfaces the cases you skipped, and forces you to answer questions you would otherwise have left implicit. Those are the questions that normally get discovered three weeks later, in a demo.

The loop is: deep research and state of the art, then lay out the ideas and the problems, then draft the spec with the agent, then read it back critically. It is long, and a large share of the value sits there.

One prompt does more work than any other. Append it to every feature request:

Before you answer, tell me what you need to know to answer well, and point out any assumptions you'd otherwise make.

It converts silent bad assumptions into explicit questions, which is the failure mode that costs weeks.

A practical asymmetry appears once the framework is in place: the WHAT gets constantly challenged, the HOW less and less. The spec determines the functionality and the tests, so that is where the effort belongs, and where a bad decision is expensive because everything downstream inherits it. The architecture gets challenged less over time, not from complacency, but because the rules are already encoded in the framework. With them in place the AI produces clean, testable architectures by default.

That asymmetry is where the framework pays for itself. Writing the rules demands real architecture skills and deep knowledge of the stacks, but they are written once and everyone benefits from them on every project. Individual developers stop re-litigating layering, error handling and testability on every feature, and spend their attention on the requirement instead.

5. The documentation model

This is the core of agentic development and the biggest single lever on result quality. Four document types, split by topic, with a global index, a per-folder index and a code map.

DocumentAnswersThrilled
Specthe WHATRequirements. Source of truth. Only the WHAT, never the HOW. Ideally testable
Designthe HOWImplementation, chosen patterns, design system, stacks
ADRthe WHYWhy a decision was made, under which constraints, which limitations were knowingly accepted
Runbookthe HOW TO OPERATEProcedures around the project, not the code: deploy, back up, recover
SPEC WHAT Requirements. Source of truth. Current state only, no history. Ideally testable. DESIGN HOW Implementation, chosen patterns, design system, stacks. ADR WHY Decisions, their constraints, and the limitations knowingly accepted. Immutable. RUNBOOK OPERATE Procedures around the project, not the code: deploy, back up, recover. GLOBAL INDEX + PER-FOLDER INDEX + MAP CODE Every fresh agent starts here, and gets a short exploration phase instead of re-walking the tree. Transient specs and plans live outside this model, and are deleted once the task is done.
Four document types, one index. The split is what makes starting every task from a clean session affordable.

The split is not bureaucracy. Each document answers a different question, and mixing them is what produces documentation nobody maintains:

  • Merge WHAT and HOW, and every implementation change invalidates the requirements.
  • Lose the WHY, and six months later someone “simplifies” a constraint that existed for a reason. An agent will do it in an afternoon.
  • Keep transient plans around, and the next agent reads a stale plan as current state.
  • Let change history creep into the spec, and the agent can no longer tell what is true now from what was true before. The spec carries the current state only: no history, no “previously we did X”. History belongs in ADRs and in Git.

Why the indexes matter

The goal is to start every task from a clean session. That only pays off if a fresh agent can understand the project without burning its context rediscovering it. Index, code map and minimal docs give a short exploration phase for every new agent. Without them each fresh agent re-explores the tree, and you have paid the reset cost without getting the benefit.

Two rules go with it. No transient documents in the repository's docs: the iterative specs and plans produced during a task are verbose and bound to it, so delete them once it is done. The implementation plan in particular is never re-read and misleads more than it helps. And docs are written for agents first. They do not have to read beautifully for humans. If you have a question, ask the agent rather than digging.

The context failure mode

The central failure mode of agentic development: when the agent's context fills up, it goes off the rails. Three things happen, in increasing order of damage.

The context burps. As the window grows, the early tokens, which is exactly where the rules were loaded, carry less and less weight against thousands of tokens of recent output. Agents do compress and summarize, but they re-read their original instructions poorly, and compaction drops constraints first because constraints look like boilerplate.

Then it starts inventing. Gaps get filled with plausible reconstructions instead of checked facts, which is the visible symptom everyone complains about.

Worst, because it is silent, it stops applying the constraints the framework exists to enforce. The branch it was told to work on, the linter it was told to re-run, the document it was told to update. Nothing announces this. The work simply comes back subtly off policy, and you find out at review time or later.

That last point is the real argument for giving every agent a clean context, and it is not about saving tokens. In a small context the rules hold a large share of the model's attention. In a saturated one they are a rounding error. A sub-agent that receives only its task and the rules that bind it will respect them. The same instruction, given at hour three of a filled session, often will not.

Countermeasures, in order of impact:

  • One task, one session. Reset context between tasks.
  • Sub-agents with clean contexts. When the agent offers to execute in sub-agents, say yes almost always. The main agent drives (instructs, verifies), sub-agents do the work in isolated contexts. The exception is small, tightly coupled edits, where inline keeps coherence better.
  • Even then the main agent eventually pollutes itself and drives worse. Start a fresh session.

A project whose documentation is not organized this way makes all of it worse: the agent burns context rediscovering the codebase, so saturation arrives sooner and the rules go first. What looks like a hallucinating model is usually a saturated one, and that is preventable.

Here is the execution prompt we use:

Execute the whole plan in sub-agents with clear context. Use sub-agents adapted to the difficulty of the task. You, the main agent, are responsible for the correct execution and verification of all tasks and sub-tasks. There is no need to parallelize, execute sub-agents sequentially in your recommended order. Do not blindly acknowledge a sub-agent's work, every work must be challenged.

Sub-agents rather than inline buys isolated contexts. Sequential rather than parallel matters more than it looks: left alone, the agent fans everything out at once, which burns context faster, triggers provider overload errors much more often, and makes recovery after a quota exhaustion painful. Resuming one agent is simple, making sure ten restart correctly is not. And challenging every deliverable is not optional. Never settle for a sub-agent's own report of its own success.

6. Low-entropy engineering: why the technical defaults look like this

Nothing in the framework is arbitrary, and the reasoning matters more than the rule. Once you know why a default exists, you know when it legitimately does not apply. Every default can be deviated from, with a written justification in the project spec.

Language models work better in low-entropy environments

Our backend, CLI and tooling default is Go. Its advantage is not richness, it is its deliberately low abstraction ceiling: fewer syntactic choices, one idiomatic way to do each thing, an exhaustive standard library, rigid static typing, a single concurrency model. Generated code follows the same structural patterns every time.

In high-optionality environments the model faces a large combinatorial space: dozens of competing frameworks, disparate typing approaches, fragmented utility libraries. That excess optionality is what makes agents hallucinate hybrid architectures mixing incompatible paradigms.

AGENT RELIABILITY Go TS strict Python + strict typing Python, untyped JS, no gate LOW OPTIONALITY HIGH OPTIONALITY one idiomatic way · compiler catches it competing frameworks · hallucinated hybrid architectures
Agent reliability tracks how few ways there are to express the same thing.

Three properties matter, in this order:

  1. Tooling to constrain the agent: linting, static analysis, vulnerability scanning, build. More harness available.
  2. Compilation: a large share of agent mistakes explode at compile time instead of at runtime.
  3. Low entropy: explicit beats clever when a machine is writing it.

An honest nuance. The claim that Go has had no breaking structural change in a decade needs tempering. Generics were a paradigm shift that forked practice into pre- and post-generics code and partially fragmented the training corpus.

Why not Rust

The question comes up constantly, so here is the full answer rather than a dismissal.

Compilation time is the first reason, and it matters more here than it would elsewhere. An agent recompiles dozens of times within a single task, so build time does not add to the cycle, it multiplies. Fast iteration is most of the game with agents, and Rust is where that loop gets slow.

The second reason is that nobody here could audit the result. We have no Rust practice and little low-level practice generally. In an agentic setting that is not a statement about what we could write, it is a statement about what we could read: the agent produces the code, and a human has to be able to open it when something goes wrong at three in the morning. That human is the last link in the harness. Adding a second compiled language also means a second CI chain, a second review pool and a second sustainment perimeter. This is also why the earlier example of a team shipping components without reading the code is not a contradiction: that was an internal benchmarking tool, not a service with customers behind it, and the tier decides how much unread code is acceptable.

Third, our workloads do not ask for it. These are API services on Kubernetes, bound by I/O and by the database, at moderate throughput. Rust's real advantages are predictable latency without a garbage collector, tight memory budgets, and parsing untrusted input. Those advantages exist, they are simply not what constrains us. Where they would constrain us, a parser exposed to hostile input, a hot path genuinely limited by CPU, a component with a hard memory ceiling, Rust is the right answer and the framework says so.

Fourth, on our own criterion, Rust is the higher-optionality environment. Traits, macros, generics and a choice of async runtime give more ways to express the same thing than Go does, which is the property we deliberately optimized against.

The serious counter-argument deserves stating. Rust's compiler is the strictest harness on the market, and this article argues that types and compilers are harness, so Rust should follow. Our experience is that the compiler does catch more, but the cost per attempt is higher, and borrow checker and lifetime errors are where we see agents flail, changing signatures until something compiles rather than understanding the constraint. That is an observation from our practice, not a law, and it is the argument most likely to change.

Last and least, generated Rust suffers from ecosystem churn more than from corpus size: competing async runtimes and successive generations of error-handling idioms mean models mix conventions that do not belong together.

None of this makes Rust a bad language. It makes it the wrong default for what we build.

Python is restricted, not banned. Go has no mature ML ecosystem, and data engineering, ML pipelines and scientific processing legitimately live in Python. The rule is: if it can be done in Go, do it in Go. When Python is approved, deterministic lockfiles and reproducible environments are mandatory, and strict typing is required. Strict typing claws back some of the entropy that made Python the weaker choice for generation in the first place.

On the frontend, TypeScript with strict mode mandatory, for the same reason Go wins on the backend: types are a harness that catches agent mistakes before runtime.

A formal contract between backend and frontend

Between a Go service and its frontend, the contract is an OpenAPI specification acting as single source of truth, with generated types on both sides. The backend changes an endpoint, the spec is regenerated, a frontend CI job re-runs the generator and blocks the merge on an uncommitted diff, and components using a renamed field fail to compile before anything deploys. This kills an entire bug class, contract drift from hand-copied types, which is the kind of mistake an agent makes silently and repeatedly.

Crash early, let the orchestrator handle it

The default error model is let it crash, borrowed from Erlang/OTP and adapted to Kubernetes. Transient errors get one to three local retries with backoff. Past the limit, log an error and exit non-zero. Structural errors (missing configuration, absent dependency, inconsistent state) crash immediately, because the problem will not resolve itself.

Why crash rather than cope? Kubernetes is the recovery orchestrator, with restart policies, backoff, probes and rescheduling. Reimplementing that in application code produces complex, hard-to-test, frequently buggy logic that does the orchestrator's job worse. A service that crashes cleanly with a clear message is more reliable than one that survives in a degraded state at all costs.

This matters doubly with agents. “Handle every error gracefully” is the kind of instruction that makes a model generate elaborate defensive machinery nobody asked for.

Zero-diff linting, or linting against the machine

AI-generated code has different anti-patterns from human code: excessive defensive code, orphan imports, unsolicited recreation of existing abstractions, and security flaws hidden under syntactically perfect surfaces. Code-quality analyzes put readability problems around 3 times more frequent and formatting problems around 2.66 times more frequent than in human-written code.

Two decisions follow. The linter belongs inside the agent's loop, with an explicit instruction not to stop until the checks pass, so a failing linter is redirected to the agent as error feedback and forces self-remediation. And zero diff is allowed on formatting and static rules. The point is not style, it is removing dead code, capping complexity, catching obvious bugs and filtering security issues before a human review. CI is the last-resort guardrail, not the discovery mechanism.

TDD and dependency injection as control mechanisms

TDD inverts the model's natural failure mode. Unconstrained, generative models produce defensive, over-architected code with useless imports and abstractions anticipating improbable scenarios. Requiring a failing test before any implementation forces the agent to satisfy one verifiable constraint. It cannot wander into unrequested features.

TDD also reduces comprehension debt, because the tests are living, executable documentation of intended behavior. And it is the best answer to “I don't trust AI code”: you define the contract, the AI implements it, the tests verify. If the AI produces bad code, tests fail immediately, and there is no need to read 200 lines hunting for the bug.

Dependency injection is what makes that possible, so it stops being merely good design and becomes a control mechanism. It buys explicit interface contracts, since a constructor signature reveals the component's requirements and the agent cannot hide coupling inside business logic. It buys interface segregation, which has to be enforced because LLMs optimizes for generation fluency and drift toward monoliths. And it buys strict side-effect isolation, with network and database calls pushed to the boundaries and injected, so the agent iterates on pure logic without touching external state.

Test observable behavior, never internal implementation details, otherwise every minor refactor invalidates dozens of tests.

Spec and tests together

Specification-driven development alone shows real weaknesses on existing codebases. Each intermediate textual step amplifies hallucination, so a decision validated during requirements gets quietly altered by the model at design time, because there is no strict behavioral coupling between spec and generated code. TDD alone has the opposite gap: it verifies behavior rigorously but cannot encode business intent, non-functional constraints or architectural decisions.

The opposition is a false dilemma. Textual specs carry the WHAT and the WHY, automated tests encode the verifiable HOW. That combination prevents spec drift while preserving traceability, and it is why the documentation model separates spec from design from ADR in the first place.

The workflow adapts to the task

The canonical loop is a backbone, not a liturgy. Bug fixes skip product shaping and architecture review entirely, and they are the one category mature enough today for near-autonomous ticket-to-agent execution. Backend features get the full loop, because they are testable end to end. The judgment call stays with you: applying the heaviest workflow everywhere burns tokens and goodwill, applying the lightest one to a production backend feature is how you ship a mess.

UI work is the clearest case for dropping the process, and it earns its own mode. Running the full spec and TDD treatment to move a button is a tank for a nail: the agent writes a test, fails it, changes the color, greens the test, runs every linter, and then you look at it and want a different color. That is a lot of tokens for nothing. Worse, UX is much harder to describe to an agent than code, so it usually lands off target and the only loop that works is seeing the rendering. So we switch the process off explicitly:

Switch to UI/UX editing. Create a branch, commit all edits but do not push, and do not run checks or tests until I say that we have finished UX editing.

Then the frontend runs with hot reload, changes are visible live, and you iterate by looking rather than by specifying. When it is right, switch the process back on:

Stop UI/UX editing; make local checks and tests. If OK, then squash all commits from this UX/UI branch while keeping a detailed commit message, and push.

The checks and the tests still run, and the branch still lands as one reviewable commit. They just stop running dozens of times while you are deciding what you want. This is worth transposing to any problem where you only know what you want once you see it, which is exactly the shape of the dashboard difficulty below.

7. Scaling the constraint: third party

A single standard is either too heavy for a personal script or too light for client production. So projects declare a tier, and the constraint scales with the blast radius.

ThirdNatureConstraints
APersonal projectLight rule file, few sharing constraints. Optimized for iteration speed
BShared, internal audienceIntermediate. Documentation and changelog expected
CShared and going to productionHeavy: monitoring, log management, failure handling, dashboards, SLA

Pick honestly. Tier C on a throwaway script wastes your week, tier A on something that reaches a client is how incidents happen. The cost of tier C is real, and paying it on a disposable tool is as much a mistake as skipping it on something a customer depends on.

8. Multi-agent review

This is the quality practice with the best return, and the one that catches the rules your main agent forgot.

The protocol:

  1. The driving agent asks his questions and produces a review plan plus three independent prompts.
  2. Three reviews run in clean, mutually blind contexts, written to separate files. No reviewer sees another's output.
  3. One of them then writes the synthesis, with the instruction to validate each claim raised by the others, ending in a remediation plan.
The artifact code · spec · plan Reviewer A clean context, own file Reviewer B clean context, own file Reviewer C clean context, own file no reviewer sees another's output Synthesis validate every claim Remediation plan
Multi-agent review. The value is not in any one reviewer, it is in the blindness between them.

For pure code review the major model families are equivalent. Each finds shared issues and independent, relevant ones. The point is not picking the best, it is crossing them. Even single-provider crossing works: a mid-tier and a top-tier model of the same family already surface different things.

Two observations from practice. On code that has been reworked several times, reviewers often find nothing, which is a good sign and a useful maturity signal. On recent, less mature areas, typically deployment code, each of them surfaces different things.

A concrete example of what it buys. On a monitoring CLI, the main agent defaulted to logging on stdout, following the service rule it had learned. All three independent reviewers flagged the violation: a CLI runs in a terminal, not a container, so stdout logging pollutes the output, and it should log to a file or syslog instead. The main agent had not seen it itself.

Cost framing matters here. Agentic usage consumes 5 to 20 times the tokens of plain completions. A cross-audit by three or four frontier models over a significant codebase is not free. Use it for critical components and major architectural decisions, not for every routine merge request. A single second-model pass on a structural change gives roughly 80% of the benefit for 20% of the cost.

On economics generally, one platform built over five weeks would have cost well over $10,000 in tokens billed per API call, versus a handful of flat-rate subscriptions actually used over that period. Pay-per-use API pricing is not manageable for daily agent work.

9. What actually goes wrong

Each of these costs someone real time.

Do not stack up tasks. Early on a large project, results came so fast that a whole batch was handed over at once, with automatic relaunch at every quota reset so it would keep going unsupervised overnight. Roughly half of what was produced had to be thrown away, plus a large amount of time spent unpicking the off-topic parts. One subject at a time, verify, consolidate the docs, clear the context, next subject. “Figure the whole thing out on your own” does not work on a complex project.

Never give a demo mockup as a frontend spec. An interactive mockup containing fake JavaScript was handed over as input for coding the frontend. The agent mixed the demo code into the real code, and it nearly cost the whole codebase. Give the visual plus a description of the steps and interactions, advance piece by piece, never the demo code.

Never trust the agent's self-reported verification, especially on the frontend. On a simple frontend feature, an agent was explicitly asked to verify visually with a browser automation tool. He claimed it had, and that the result matched. Shown an actual screenshot, it admitted the result was nothing like what had been asked, and even after a second attempt it still was not right. On the frontend, agents do somewhat what they want despite clear specs. On the backend (RPC, message queues, mutual TLS, token handling) the same team reports essentially no problems, because those technologies are testable. Check the rendering yourself.

Rewriting an existing project: never ask for a line-by-line translation. Three attempts at porting an internal CLI from Python and shell to Go:

  1. “Rewrite this script in Go, iso-functional, go ahead”: total failure. The agent never understood how it worked. All thrown away.
  2. Feeding it a one-hour transcript of the original author explaining the project: slightly better, still far short. Even the author had not managed to explain everything in an hour.
  3. What worked: drop the idea of rewriting. Explore the code together with the agent, confronting your own understanding against it, until you have a spec of the actual need, discarding the features that only existed because of the original author's particular design. Once that spec was clear, developing from scratch was straightforward.

The agent does not always know where it is pointing. A request to test a new component on a pre-production pipeline led to two chained mistakes. The agent asked for a message to be pasted into a message-queue UI, it was done without thinking and in the wrong UI, so the message went to production. Then the agent pushed its own test messages missing a required identifier, which blocked the connector feeding the search cluster. Nearly the entire infrastructure of that team went down. Read what you execute, and state the target environment, and what is production, explicitly to the agent. Note the inverse observation elsewhere: on other projects the agent asks for confirmation systematically, even on pre-production. Behavior depends on the framing, so never rely on it.

“Go ahead, figure it out” can go very far. Asked to fix a firewall problem on a personal project, an agent noticed SSH was open with the user's key, connected to the machine, pulled the scripts and started operating on production. It went fine. It could have gone very badly.

Check the pipeline is entirely green. The agent often looks at the first job, sees green and concludes it is done while later jobs fail. Ask it to iterate until fully green, and expect it to skip local checks before pushing despite explicit rules. CI is the real guardrail.

Where agents shine, and where they do not

They work well on standalone, well-framed projects with defined inputs and outputs. A team with no ML background brainstormed the subject with an agent, which then built a local test and benchmark platform, ran whole nights of model tuning, and produced two qualification components that work well, without anyone reading the code. Also: infrastructure-as-code reviews, sysadmin companionship, capitalizing operations into the docs, investigating live application logs, driving version-control CLIs.

Dashboards are the contested case, and reports differ enough between teams that we do not treat this as settled. Some get usable first drafts, others spend more time explaining what they want than they would spend building it by hand. The split seems to follow the data more than the tool: already-curated metrics give the agent something to reason about, while raw log data needs meaning assigned to it before a dashboard means anything. Where it does go badly, the diagnosis is definition rather than execution, because with a dashboard you discover which indicator is missing while building it, and the round-trip with the agent breaks that loop. Two leads: ask the teams getting good drafts what they do differently, and transpose the live-rendering workflow that works for UI.

10. Domain experts come inside the loop

This is one of the largest wins, and it forces a question about roles that deserve to be asked out loud. The chain that disappears is one a product manager often stood in the middle of, translating business intent into something a developer could act on. That translation function is the part that loses its reason to exist. What does not disappear, and becomes more valuable, is the rest of the job: arbitrating between demands that cannot all be satisfied, sequencing, holding the product line, defining what success would look like, and saying no. The mistake would be to read this as the role becoming redundant. It is the intermediate half that is.

The classic way a feature went wrong had nothing to do with technology. A domain expert tried to put a need into words. An intermediate translated it. A developer specified it in their own terms. Someone implemented it. Weeks later a demo revealed the result was not really what was wanted, and nobody had lied at any step. The need had simply been re-encoded four times, losing a little at each hop.

Agentic development collapses that chain. The person who holds the need can now produce something concrete, a mockup, a working prototype, a rough script that does the real thing on realistic data, without waiting for a developer to be available and without learning to code first.

BEFORE: FOUR RE-ENCODINGS Domain expert Intermediary Developer Implementation Demo, weeks later − − − − a little of the need is lost at every hop NOW: ONE LOOP Domain expert holds the need, judges the result Agent challenges, drafts, builds a mockup people can react to, in a day
The translation chain collapses. Nobody lied at any step. The need was simply re-encoded four times.

That changes what a domain expert is expected to produce. A mockup or prototype is worth more than a paragraph of requirements, because it is something people can react to. “Not that, more like this” is a more reliable signal than any specification review. They are also the best person to write the spec's WHAT, since they know what correct looks like on real data, and the best person to verify functionally, since judging whether the output is right requires no ability to read the code.

The risk: shadow development

Without a frame, this produces shadow development, the technical equivalent of shadow IT: unversioned, untested, unmaintained code accumulating invisible technical debt until it becomes critical. The goal is not to forbid it, it is to channel the creative energy into safe rails.

Four principles, and a classification.

  1. Approved tools only. No unvalidated tool keys internal or sensitive data. In a cybersecurity context this is not negotiable.
  2. All code lives in version control, with a minimal CI. Deep Git mastery is not required. Branch, merge request, automated review is enough.
  3. Mandatory review before deployment. Non-developer contributions are treated as drafts: the AI produced a first pass, a developer validates architecture, security and integration.
  4. Templates and starters. A pre-configured project (structure, linters, CI, rule file) so the agent works inside a constrained frame rather than from a blank page. The rules apply to AI-generated code whoever prompts it.
CategoryExamplesGovernance
ExplorationUX prototypes, throwaway POCs, one-off analyzesVersion control optional. No deployment. Internal use only
Internal toolingAutomation scripts, dashboards, log parsersVersion control mandatory. Automatic linting. Developer review before production use
Production codePlatform components, business rules, APIsFull workflow: version control, TDD, CI, review, squash merge

Be honest about which row you are in, and note that things move down the table over time. The script “just for me” that a colleague starts relying on has become internal tooling.

Two cautions apply specifically to this path. A mockup is an input for discussion, never an input for code, as the war story above shows. And the data rules apply identically to everyone: no secrets, no client data, anonymize first.

11. Adopting this in a team

The condensed playbook:

  1. Start from the chore they hate, not from the tool. Regression tests, boilerplate reviews, endpoint documentation. The message is not “trust me”, it is “look at this diff”.
  2. Make the rule files the tangible proof. Run the same task without a project rule file, with its hallucinated conventions and reinvented components, then with it. Context is the strongest trust lever there is.
  3. TDD as a psychological safety net. The agent cannot cheat if it must pass tests you wrote. You define the contract, the AI implements it, the tests verify.
  4. Pilot on a real brownfield task, not a greenfield demo. The disappointment came from existing codebases, so that is where the proof has to land. Pick something the skeptic knows well.
  5. Quantify fast. Two or three light metrics before the pilot: time per merge request, review comments, coverage. Move the conversation from “AI doesn’t work” to something observable.
  6. Respect the expertise. Never “the AI will code for you”, rather “the AI codes under your orders, and your value moves from writing syntax to designing constraints and auditing systems”.

A workable ramp for a small team: weeks 1-2, silent foundations, meaning rule files and zero-diff linting. Weeks 3-4, one volunteer on one real brownfield task, documented and shared. Weeks 5-8, a second use case, each person picking their own. Month 3 onwards, standardize what worked and drop what did not, without guilt.

One prompt worth reusing for unattended runs, overnight or while you are in a meeting:

Do everything you can without asking me questions. If you have any doubt that is not blocking, note it in a file, and we'll review it when I'm back.

12. What is not (yet) solved

Writing code by hand has become obsolete for us. In a year of practice, one hand-written equation on a personal project, because explaining it was slower than typing it.

Deployment is different. It still demands substantial expertise and real knowledge of the stacks to steer agents correctly. The framework keeps improving on that front, but it is not solved. Budget human expertise there, and expect a multi-agent review to keep finding things in deployment code long after it finds nothing in application code.

A closing deposit that has nothing to do with AI capability. Internal development targets our own efficiency and innovation. It is not meant to rebuild what specialized vendors already do well. Build internally when the need is too specific to how you work, or when a market tool would create excessive vendor lock-in, and arbitrate anything substantial.

Development being easy is not a reason to develop anything and everything.

Sources

Written from a year of building software with agents. The framework described here is internal, the reasoning is not.

Written by Stany MARCEL with Claude, from Stany MARCEL's research work and the agentic development framework built by Intrinsec's Software Engineering team.

Articles by category