Where the time goes when one person runs the products: 698 commits in five weeks
An agent writes the code to my architecture, and it writes fast. So “how fast do you write code” has stopped being a useful question for me. The better one is where the time goes once code is no longer the bottleneck. To answer it from the record rather than from memory, I took the commit history from 20 August to 25 September and labelled every commit by what it was for.
What was counted
Five products, all of them already live with users, plus this site. Only commits on the main branches, merges excluded.
| Product | What it is | Commits |
|---|---|---|
| bazito | an aggregator of Telegram classifieds across 11 countries | 317 |
| AstraBrain | AI astrology on real ephemerides, web and Telegram, 10 languages | 154 |
| TZ Portal | a public-procurement spec builder and procurement catalogs | 104 |
| Aviozo | a travel map for couples and groups: visas, flights, last-minute deals | 90 |
| Signage CMS | a CMS for a two-screen museum installation | 20 |
| this site | writing and case pages | 13 |
That is 698 commits over 37 days. Over the same period there were 92 cloud agent sessions, and that is not all of the work: some of it ran locally.
The labelling is manual, one label per commit, by its main purpose:
- data: making the product say true things. Price parsing, categories, geography, reference data, dates;
- search and growth: indexing, redirects, analytics, the funnel, translations and landing pages for search, external articles;
- operations: reliability, memory, cost, database backups, watchdogs, builds and deploys, incident work;
- features and UI: anything new for the user, and fixes to screens;
- records: handover documents, reports, the list of open work;
- security and legal: RLS, rate limits, blocking crawlers, legal pages.
| Kind of work | Commits | Share |
|---|---|---|
| Data | 173 | 24.8% |
| Search and growth | 165 | 23.6% |
| Operations | 137 | 19.6% |
| Features and UI | 94 | 13.5% |
| Records | 80 | 11.5% |
| Security and legal | 49 | 7.0% |
Two caveats. A commit is not an hour: a one-line fix and a two-day investigation weigh the same. And the line between growth and features is blurry: a landing page with a live calculator went to growth. Give features half of growth and they rise to just over a quarter, no more.
Each product leans its own way. bazito is 46% data fixes. AstraBrain and TZ Portal put about 40% into search. Aviozo has the most operations work, 28%. Only Signage CMS is more than half features, because the period went into building scenes for a new installation: 13 commits of 20.
Data: the product has to say true things
This is the largest share. In bazito it took 145 of 317 commits: price, period, currency, category, city. How each of those fixes is checked I covered in the piece on calibrating price parsing. The point here is different: the stream does not end. People write the listings, and every new city, language and habit of writing needs its own fix, from the Kazakh “пәтер” (flat) to the Montenegrin “u Budvi” (in Budva).
In the other products data breaks more quietly. In AstraBrain some celebrity pages carried wrong birth dates, which means a wrong Sun sign for a living person. Where a date could be found and checked, it was written down, 67 of them in one pass. Those pages stay closed until the reading is rewritten for the correct chart. Where the source was weak, the page was taken out of the index rather than recomputed on a doubtful date.
Aviozo ran into a similar question on its visa pages: is the date on the page the date the data was loaded or the date a person checked it? Those are different things, and the markup for search engines must not present one as the other.
Operations: failures that stay silent
The most expensive failures of the period did not put a single error on a screen.
In bazito the chat access passes moved into the database, and the function that resolves a chat started returning objects of two different shapes. Collection failed on its first line, before reading a single message. In the log it looked like a one-off error in one channel. In fact there were 18,172 such lines a day from 77 chats. For four days 126 of 179 chats delivered no messages at all, and the inflow fell roughly fourfold. The better the new cache worked, the more chats it switched off.
In AstraBrain on 11 September the balance with the model provider ran out and generation stopped. The horoscopes for the 11th were rerun by hand, and the 10th could no longer be caught up.
In TZ Portal, over seven night hours on 13 September, the code pages received 5,464 requests to 5,253 different URLs, two thirds of them from the Semrush and Ahrefs SEO crawlers. Every new URL means a render and a cache write on Vercel, so the site owner pays for that crawl. The crawlers were blocked in robots.txt and refused at nginx, before a request reaches Vercel.
That is why part of operations is watchdogs rather than fixes: the bazito pipeline watchdog messages the owner when the cycle stalls, and the nightly database copies of bazito and TZ Portal are shipped off the server.
Security breaks what used to work
49 commits, 7%. The smallest share, but two stories from the period show well why such changes cannot be rolled out blind.
In AstraBrain row-level security was switched on for every table. The payment waiting screen learned about a payment through a Supabase Realtime subscription to the payments table, and Realtime obeys RLS. The subscription went quiet, and a person who had already paid was left on the waiting screen. Access was still granted; only the screen’s automation broke. The choice was to open the table of amounts and purchase history to the anonymous role, or to remove the reason. The wait moved to polling our own route, which returns nothing but a flag. It also turned out the subscription had been worse than polling: it fired on the underpayment record too, and closed the screen as if the payment had succeeded.
In Aviozo the “Price history” button said “not enough data yet” while the table held 1,106,635 rows. A migration had applied partially: the table was created, but disabling RLS on it was not. With RLS on and no policies the API returns a 200 and an empty list, not an error. The lesson went into the handover document: an empty response does not prove an empty table.
An agent’s conclusions are hypotheses
The most useful habit of these weeks: do not act on a written conclusion until it has been measured again. An agent writes conclusions with confidence, and so does an audit. Here are five cases from the period where a measurement overturned the record.
| What was written | What the measurement showed |
|---|---|
| The bazito rebuild loses 1,333 verticals | Nothing was lost. The server held an old version of the code that silently swallowed a parameter, and the current code yields 1,695 verticals instead of 618 |
| Mobile LCP on the TZ Portal home page is 4.6 s | 1.97 s, the median of five runs. Under the 2.5 s threshold, so no day was spent on optimisation |
| The TZ Portal landing page has a 70% bounce rate | 66 of 103 visits were direct hits from bots in the US. The bounce rate from Russia is 18% |
| AstraBrain does not need a geo-based locale | It does. English is the browser default for people who do not use it, and Russian visitors were sent to /en |
| 1,248 Aviozo visa pages were checked by a person | 18 were. 1,232 carry a link from a robot, and a load date must not pass as a check |
In AstraBrain the wrong conclusion was kept in the records with its analysis, not deleted. The cost of a conclusion like that is the work that later does not get done because of it. In the empty-table story at Aviozo, hypotheses were built three times on the assumption that the table was visible as it really was. What settled it was a count returned by the writer of the data itself.
This is my answer to the question people ask about AI-written code: do I look at what it does. I do not read every line. I look at measurements: a control sample that must not change, a count from inside the process, the median of several runs.
Records: sessions do not remember each other
80 commits, 11.5%, are documents: the list of open work, handover documents, instructions for the next session. Agent sessions run in parallel and start from zero, and they share one memory: the repository.
Two of the five stories in the table above come down to exactly this. Another session shipped the geo locale in AstraBrain while the records still said it was not needed. A parallel session was fixing the visa page markup in Aviozo on the same day. A team solves this with a stand-up; for me that role is played by a written state that is checked against the code.
What follows from this
For someone hiring: the speed of writing code is no longer the signal. The signal is how many checks stand between the code and the user, and who puts them there.
For a founder with an MVP: 4–6 weeks is the build up to release. These five weeks fell almost entirely after release, because all five products were already live. New features took a seventh of the commits. The rest went into keeping a working product truthful, loud when it fails, findable in search, and closed where it should be closed. That work belongs in the plan from day one, not after the first incident.
Labelled by hand from the commit messages on the main branches of five products and this site, from 20 August to 25 September 2026, one label per commit. Most commits are made by an agent in my sessions; I set the task and accept the result. The numbers in the examples come from the commit messages themselves. Server addresses, keys and channel names are left out.