Notes from practice

The model did the producing. Catching what it got wrong was the job.

7 min readHalil Eren ÇelikProcess · Tools

A designer's worth today lies in knowing why the model decided what it decided, and in being able to think about what to do with that — to understand it, and to direct it. That is easy to say in one sentence. Building a product end to end that way teaches you what it means.

bugreports is a personal product experiment: a civic reporting system where a problem is filed with its evidence, the community backs it, the institution works it from a panel of its own, and the person who filed it confirms the fix. Strategy, product definition, the design system, the mobile app, the institution panel — all of it made with AI. What follows is the record: what actually worked, and where a person had to step in.

A library first

Prompting freehand produces a different standard every time. Ask for the same work twice and two answers come back at two depths, and the only way to see what each one skipped is to read both from the top. The fix is not better prompts. It is writing the method down.

That is what my own skill library is for: DEINS — 15 chapters, 119 services. Each service has a written definition: when it applies, the steps it runs, what it delivers, when it counts as done, and which service it hands its output to. Calling a service is calling a method, not a prompt.

The library was not built for this project; this project is what tested it. On bugreports, 93 of the 119 services ran. Twenty carried the work:

  • Experience Strategy
  • Market Research
  • Competitive Intelligence
  • Product Definition
  • Roadmapping
  • Information Architecture
  • UX Design
  • UI Design
  • Interaction Design
  • UX Writing
  • Design Tokens
  • Component Systems
  • Accessibility
  • Solution Architecture
  • Mobile Development
  • Technical Delivery
  • MVP Development
  • Design-to-Code
  • Launch
  • Quality Engineering

The division of labour: strategy, product definition and competitor analysis with ChatGPT; the design system, the interfaces, the mobile app and the institution panel with Claude Code. At every stage a skill was called, its report read, and the work approved or sent back. Nothing here is autonomous — a person stands at every step.

Approving is work too

AI does not always produce a flawless result, and not always down to the pixel. We trust it with a great deal, and most of the time we never notice that it has slipped. Being expert enough to catch that is what makes a person the key player.

The clearest example came off the project's own page. I wanted to mark the 93 services that ran. They were marked in one colour, and what came out was a list that looked authoritative and said nothing: three quarters of the page was emphasised, which means none of it was. The output was not broken. It was plausible — and plausible is the hardest thing to catch.

The fix was three weights: the twenty that carried the project in ember, the seventy-three that also ran with a light edge, the twenty-six that did not stay faint. Same data, three times as legible. The model did not make that call. It was told to mark, and it marked.

Anyone can see a broken output. Seeing a plausible one that is wrong takes knowing the work.

Product decisions are still product decisions

There is one line in the institution panel that matters more than the rest: “Resolved” is not the institution's to declare. The institution does the work and proposes closing it; the person who filed the report confirms. That is not an interface preference, it is the spine of the product — leaving the resolution rate to the institution's own word would have removed the reason the product exists.

No model volunteers a decision like that. A model produces the actions — approve, reject, mark resolved. Who holds which one is said by the person thinking about the product.

Two demos that actually run

Neither demo on the site is a screenshot; both are the compiled applications themselves. The mobile app is an Expo web export: three megabytes of JavaScript that never calls a server and carries no keys of any kind. The institution panel is a static export of a Next.js App Router build. Both really run in the browser.

Getting there produced the work that hides under the sentence “AI wrote the product”:

  • The fonts 404'd and the app never mounted. Expo writes bundled dependencies under a folder whose name begins with a dot, and Vercel will not serve a path segment that starts with one. The paths were flattened and the bundle rewritten.
  • The photographs came from a third party. They were localised — so the demo does not go dark when that site is slow, and because the capture rig has no route to it.
  • The phone's safe area came back as zero. A browser reports an iPhone's notch and home-bar insets as nothing, so the app's header ran into the island and its tabs into the bar. The real values are handed to the frame as a parameter.
  • The panel's fields looked like trenches. It was carrying a far deeper inset shadow than the mobile app's own — at tablet size it read as a ditch around every field. Both demos were brought onto the same token values.
  • Fonts were being fetched at build time. Next.js's Google-font plugin goes to the network during the build, and the build machine had no route there. The fonts moved into the repository and are served by plain rules.

None of this is design genius or engineering brilliance. It is the residue of making something actually run. And a model does not hand it to you: it produces the bundle, and you are the one who opens it to see whether it works.

The stutter

This was the most instructive failure. The demo stuck under a trackpad: the finger kept moving, the list did not. My first hypothesis was wrong — I assumed the machine was struggling, that frames were being dropped somewhere.

So I measured. Frame timings were recorded across the scrolling: of more than two thousand frames, two went over 33 milliseconds and there was a single long task. Rendering was never the problem.

Then the events. Wheel events were reaching the app; not one of three hundred and sixty-three had been cancelled. No code was blocking them. The browser was scrolling — just not the list under the pointer.

Then elimination, one thing at a time. The page lock was disabled: no change. The frame's scale transform was removed: stalls fell from 73% to 42% — a factor, not the cause. Next in line was a single CSS rule I had put inside the demo document myself.

That rule was there to solve a real problem: with the pointer over the phone, a wheel step the app could not use chained out of the frame and scrolled the page underneath it. The fix had switched overscroll off on every element. Removing it took the stalls from 73% down to nothing worth measuring — and the page still did not move a pixel.

The mechanism is this: a wheel gesture latches to the scroller under the pointer the moment it starts. With overscroll contained on every element, a chip row, an avatar strip or a short inner list swallowed the rest of the gesture as soon as it hit its own end, instead of handing it up to the outer list. Keeping the rule at the root alone holds the page still and lets the list scroll.

An older fix for the same problem had been worse: a non-passive wheel listener. That listener forces the browser to run JavaScript before it may scroll, which takes the frame off the compositor — on a trackpad the demo scrolled a frame behind the finger.

No model would have found this. Not because it is hard, but because it only exists in the hand.

The record of this failure only begins when a person moves a trackpad and says something is off. The measurement is what turns that feeling into a cause. Both of them stayed with me.

An invented world

The demos had to look real without carrying anything real. The decision: no real institution, brand or place appears in either of them. The names read as real, so the demo stays believable — and no institution is shown a resolution rate it never earned. Two hundred and forty-five replacements: the city, the districts, the transport authority, the water authority, the brands, all of it.

That is another decision a model does not volunteer. What it volunteered was believable sample data. Seeing where believable becomes a problem is separate work.

So who wins

AI's great gift is that it hands you more speed and more capacity than you asked for. That is the reason one person can carry a product from strategy to two running demos — what used to need a team and a budget is now possible with one person and a method set up properly.

But what is handed over in abundance is speed and capacity, not judgement. Judgement still sits with whoever knows where to stop, what looks plausible and is wrong, and which decision belongs to whom.

As long as we put that to proper use and position ourselves accordingly, we are the ones who come out ahead.

bugreports is a running demo: no users, no pilot, no institution, and every figure is sample data. The project page opens the demo in the phone and in the panel.