When agents write the code, an FDE's work comes down to three things: definition, constraints and judgement
The faster agents write code, the clearer it becomes which work still needs people: understanding the problem correctly, setting limits for the agent and saying plainly which results are worth anything.
In brief
- GitHub analysed more than 2,500 agent configuration files and found most failed because they were too vague. The hard part is specification, not code generation.
- Cognition argues that faster code generation only pushes the problem on to testing, review, deployment and maintenance. Enterprises want value, not a system that merely looks busy.
- The three jobs FDEs keep are naming the user and the success criteria, writing the spec as a living document, and judging results by understanding the code rather than waiting for tests to pass.
According to Addy Osmani, GitHub analysed more than 2,500 agent configuration files and found that most of them failed for a mundane reason: they were too vague. The people who wrote them had not said clearly what they wanted.
That finding matches what Cognition presented at the AI Engineer World’s Fair 2026. In the company’s view, generating code faster does not solve the whole problem, because testing, review, deployment and maintenance are all still there. The bottleneck has not gone away. It has moved.
Developers hoping to move into FDE work should take note. Typing code quickly is no longer the skill people pay you for. The valuable part is the three jobs agents cannot yet do for you: defining the problem, constraining the agent and judging the result.
Freelancers execute, FDEs decide
In his essay on forward deployed AI engineering, Ghanemzadeh draws a neat line: freelancers execute, FDEs decide. As he describes it, an FDE is an operator with opinions about what to build, in what order, how to build it and how success will be measured.
The first question of definition sounds simple, yet it is the one most often skipped: who is the user? Ghanemzadeh puts it bluntly. If you cannot name the user, you do not have an FDE engagement. A request such as “the client wants an AI agent” is a wish. It is not yet a problem.
Picture a bank that wants an agent to process loan applications. “Automate underwriting” sounds specific but gives no direction.
Change it to “branch credit officers, who spend most of each day cross-checking documents”, and you immediately know what to measure, what to tackle first and when the work counts as done.
An agent cannot take this step on its own. It does not sit beside the credit officer or notice which tasks frustrate them. That is why definition still belongs to the person on site.
Constraints are a living spec, not a prompt
A post on the fde.academy blog describes the engineer’s new role when working with agents as defining the task, setting constraints and then steering the agent. This is a practitioner’s view, not research.
Even so, it fits the GitHub finding above. The agent configuration files failed because they were vague, which means they failed at the constraint stage.
Osmani cites GitHub’s view of spec-driven development, in which the spec becomes the shared source of truth: a living, executable document that changes as the project does.
The word “living” matters. A spec is not something written once at kickoff and filed away. It is a contract between you, the customer and the agent, updated whenever the field teaches you something new.
Go back to the bank. A good constraint might be: “the agent only flags applications with missing documents and never rejects an application by itself.” That one sentence does two things. It limits the agent’s behaviour, and it records a business decision that anyone writing code later must respect.
Writing that sentence requires you to understand the customer’s risks, not just prompt syntax. Constraints therefore connect directly to definition: if you do not know who the user is, you cannot know what must never happen to them.
Why has judgement become the new bottleneck?
If definition and constraints are the inputs, judgement is where the pressure builds most. Osmani names the problem directly: AI generates code far faster than humans can evaluate it. Review has become a throughput problem.
Many teams’ first reflex is to lean on tests. Osmani agrees tests are necessary but says plainly that they are not enough. Tests check only what their authors thought of. The question “does this code do what the customer needs?” still has to be answered by someone who understands it.
He goes further. The familiar organisational assumption that reviewed code is understood code no longer holds. When one person approves dozens of agent-written PRs a day, the approve button confirms only that the code was looked at. It does not prove anyone actually understood it.
Cognition takes the argument to the enterprise level. Large companies in heavily regulated industries want to know whether a system creates value, not whether it is busy. An agent can run all night, open hundreds of PRs and still deliver nothing.
Telling “busy” apart from “valuable” is a judgement, and that judgement rests on the definition of success set at the start. The fde.academy post concludes that the judgement layer does not get automated just because the code-writing layer has sped up. That is an opinion, but the evidence from Osmani and Cognition leans its way.
Three jobs, three ways to fail
Looking at all three jobs together shows that each has its own failure mode, and any of them can be hidden by an agent that is working very hard.
| FDE job | Question to answer | Sign it is going wrong | How to practise |
|---|---|---|---|
| Definition | Who is the user, what gets built first, how is success measured? | The user cannot be named; the request is just “build an AI agent” | Write one sentence describing the user and one success metric for every ticket |
| Constraints | What is the agent allowed to do, and what must it never do? | Vague config files or specs; a spec written once and abandoned | Maintain the spec as a living document, updated after each agent mistake |
| Judgement | Does the result create value, and do I understand the code? | Treating passing tests as done; treating reviewed as understood; mistaking “busy” for “valuable” | Explain every change yourself before approving; measure results against the criteria you defined |
Judgement feeds back into definition
Cognition describes delivery to customers and product feedback as a single loop, in which what is learned in the field shapes product decisions. For these three jobs the loop has a clear shape: judgement at the end sends information back to definition and constraints at the start.
A useful diagnostic when judgement finds that results are not delivering value: before blaming the agent, check whether the user was named correctly and whether the spec is still vague anywhere. Fixing the input is usually cheaper than fixing the output.
This is also why the three jobs are hard to split across three people. The person who defines the user should be the one who judges the result, otherwise nobody notices when the spec has drifted from reality. That is why FDE is a role, not an assembly line.
Demand for the role is rising. The Pragmatic Engineer reports very strong hiring demand for FDEs at Google, OpenAI and Anthropic. When the companies building agents are themselves hiring people to stand between agents and customers, it is not hard to guess which skills they are short of.
What should Vietnamese developers practise?
Many Vietnamese developers grew up in outsourcing, where requirements are fixed by the client and the team simply executes. That is exactly the model Ghanemzadeh calls the freelancer. To move into FDE work, you have to learn to ask the questions that someone else used to answer for you.
Start with your current job. Before handing a ticket to an agent, write one sentence naming the user, one measurable success criterion and three constraints the agent must not break. If you cannot, the ticket is not ready, whether for an agent or a person.
When reading FDE job descriptions, look for phrases such as working directly with customers, defining requirements and owning outcomes.
On your CV, do not just list the features you built. Explain how you identified the user, which success criteria you set and how you verified that the system created value rather than merely ran.
Agents will write code faster still. Recruiters will not ask how many lines of code you can generate. They will ask whether you know which lines are worth keeping.
Was this article useful?
Thanks for the feedback!
6 sources
- How Forward Deployed Engineering is done at Cognition
- Why AI Agents Are Changing the Forward Deployed Engineer Role · 2026-07-23
- What Forward Deployed AI Engineering Actually Is · 2026-06-24
- How to write a good spec for AI agents · 2026-01-13
- Comprehension Debt - the hidden cost of AI generated code. · 2026-03-14
- The Pulse: Forward deployed engineering heats up again · 2026-05-14