Guide
The feeds are flooded with the "app in 20 minutes" video. Nobody films the three weeks that come after. That gap is, basically, my job.
By Miguel Alejandro Hayes· Founder, Hayes Projects

Every other day someone sends me a video. A guy types one sentence, the AI spits out a whole application, and something that works shows up on screen. The message with it is almost always the same: "so why do I need you?".
It is a fair question. And the honest answer is not "because the video lies". The video does not lie: the app exists, it runs, it does what it promises. The answer is more uncomfortable: what you see is the cheap 20% of the work, and the expensive 80% starts the moment the creator hits "stop recording".
I have had this argument with myself as much as with clients, so I stepped away from the tools' marketing and the testimonials of the founders who sell them. I went to the controlled trials, the papers, the large surveys. What I found is that AI is real and powerful, but almost everything you are told about "speed" and "three-person teams" is measured in a way that inflates the number.
In 2025, METR — a group that evaluates models, not one that sells them — ran a real randomized trial. They took 16 developers with genuine experience, people who had spent years in their own large repositories, over a million lines. Real tasks of a couple of hours, and a coin flip task by task: AI allowed or AI banned.
With AI, those tasks took 19% longer. Slower, plain and simple.
The juicy part came next. They asked those same developers how much time they thought they had saved. They said 20%. They were convinced they had gone faster while they were going slower: almost forty points between what they felt and what the clock recorded.
Most of the happy productivity numbers going around are just self-perception.
On top of that, METR repeated the experiment in late 2025 and the sign seems to have flipped: maybe now it helps a little. But they themselves call it "very weak evidence", because the most enthusiastic developers refuse to work half the time without AI and drop out of the sample. When the treatment becomes valuable, the experiment breaks. In short: it probably helps somewhat now, but nobody knows how much, and anyone who gives you an exact number is selling you something.
In 2023, a GitHub and Microsoft experiment measured people implementing an HTTP server in JavaScript, from scratch, as fast as possible. With Copilot they finished 55.8% sooner. That is where the "AI makes you twice as fast" you see everywhere comes from.
Look at the conditions: a textbook task, no legacy code, with thousands of identical examples in the training data. It is the ideal scenario. It is that same 20-minute video.
So AI performs best where the task is new, bounded and standard, and worse — sometimes negative — where the code is mature and you already knew it better than it did. It substitutes explicit knowledge (syntax, APIs, common patterns), which is exactly what a junior lacks; it does not substitute the tacit knowledge of a system with years behind it, which is what a senior has. That is why the junior gains a lot and the expert in their own code sometimes loses.
In the middle — a real company with a mix of tasks — the honest number is more modest: the largest field study, with nearly 5,000 developers, found 26% more pull requests completed. It sounds good, and it is. But that measures volume, not value. More PRs is not the same as better software.
This is where I live.
A demo optimizes the "happy path": the user does exactly what they should and everything goes well. Production is the other thing, the part you do not see: authentication, permissions, the user who dumps garbage into the form, the one who tries to see data that is not theirs, deployment, monitoring, and maintenance for the next two years.
How much does that gap cost? There is a brutal figure. In March 2025, a researcher scanned the apps in Lovable's showcase — one of these "app in minutes" tools. Of 1,645 projects, 170 were leaking data: names, phone numbers, payment information, API keys. One in ten "finished" applications left the door open. And it was not a sophisticated hack: it was misconfigured database permissions, that boring work the demo skips.
It is not an isolated case, it is structural. Independent analyses flagged security problems in roughly a third of the code Copilot generates. Another study found that duplicated code multiplied by eight in 2024: AI would rather copy and paste than reuse well, because copying is what it knows how to do. And at the organization level, Google's DORA report, with thousands of professionals, found the thing that sums it all up: AI raises throughput and, at the same time, raises instability. You ship more, and more breaks.
AI is (only) an amplifier: it multiplies what you already have, for better or worse.
If your team has discipline, tests and a solid platform, AI multiplies that strength. If it is a mess, it multiplies the mess, and faster. The tool does not give judgment to whoever lacks it.
The gap between demo and production does not shrink with AI; it grows.
That is why the nuance matters: AI cuts time-to-demo enormously and time-to-production far less. An "app in 20 minutes" video can be true and still correspond to weeks of real work before anyone can use it safely.
This is the part where the marketing becomes almost comical. They show you companies with two million dollars of revenue per employee and tiny teams. What they leave out is that those companies sell AI as the product: their revenue scales with machines, not with people, and they are picked for the headline precisely because they did well. It is pure survivorship bias; you do not see the thousands of startups that failed. That does not measure how much a team that simply codes with AI shrinks: it measures a different business model.
When someone actually sits down to measure it seriously — a Harvard working paper on AI-native firms — the real number is that they are between 12% and 25% smaller. Not nothing, but a long way from the "6× fewer people" of the pitch.
Where there is a solid and worrying signal is in junior hiring. Payroll data from Stanford shows that employment of developers aged 22 to 25 is around 19% below where it should be, and the adjustment comes through less hiring, not layoffs. The logic is simple: if AI covers the junior's explicit work, fewer juniors get hired.
Tomorrow's seniors come from today's juniors. If we do not train juniors, in ten years seniors will be missing.
Although we have to be honest: part of that drop started before ChatGPT, with the hangover from pandemic over-hiring and interest rates. Not all of it is AI's fault.
The cost of AI per developer is not the constraint. According to Anthropic's own documentation, a developer with an agent spends around 13 dollars per active day: about eight minutes of skilled work. If it saves you more than eight minutes a day, it already paid for itself.
The trap is in the tail, not the average. The cost explodes when you launch many agents in parallel or work on huge codebases, because — a revealing figure — most of the tokens go into reading code, not writing it: around 76%. The agent spends its time hunting for where things are. That is why a large, messy codebase is both more expensive and slower with AI at the same time.
It is not serious to say "AI changes everything" or "AI is overrated". Both are intellectual laziness. Here is what the evidence and the practice let me claim.
I use AI without guilt where it shines: prototypes, greenfield, boilerplate, tests, getting something off the ground. There the video does not lie and the gain is real.
And I hit the brakes where the evidence says it deceives: mature, critical code, where the clock can run backwards while the feeling runs forward. There the expensive work — security, permissions, edge cases, review — is not automated: it is done, with attention and detail. That is the difference between something that looks finished and something you can put in front of a client without leaking their data.
Anyone can do the demo now. The rest is the craft, and for now it is still human.
So when someone sells you speed, ask them: speed to where? To the demo, or to production?
An honest note: I wrote this with AI’s help. I used it to organize the research and tighten the prose, then reviewed, corrected and signed off on every line myself. Which is, more or less, the whole point: AI speeds up the draft — the judgment and the responsibility stay mine.
It depends on the task. On something new and bounded (a prototype, boilerplate) it can speed you up a lot. On mature, complex code, the best studies measure small gains or even time losses. The "55%" from the ads comes from an ideal task, not your real system.
You can have a demo in minutes. Production is another thing: security, permissions, edge cases, deployment and maintenance. In a real scan, one in ten "finished" apps built with these tools was leaking user data. The video does not do that work.
The serious evidence shows AI-native firms 12% to 25% smaller, not the "6×" of the marketing. The most measurable effect is less junior hiring. Translating productivity into 1:1 cuts is a mistake: demand for software grows too.
For a single developer, around 13 dollars per active day: cheap next to a salary. The cost explodes with many agents in parallel or huge codebases, where most tokens are spent reading code, not writing it.

About the author
Economist and essayist turned developer. He founded Hayes Projects, a Miami venture studio and custom software lab, to build software that ships, scales and solves real problems.
Meet the founder→If you want more than a demo — something that reaches production, secure, tested and maintainable — let's talk for 20 minutes. No pitch.
Book 20 min →Prefer email?
No obligation. No spam — just one useful, concrete reply.
We use first- and third-party cookies (Google Analytics and Google Ads) to measure traffic and improve the site. You can accept or reject them. Learn more