OBIE · POLICY PAGE · THE DELIVERY SYSTEM

It looks like a prototype. It's built like a product.

On a Saturday morning in August, I split one master brief into six build plans, assigned each to its own governed AI agent session, and took my dog for a 45-minute walk. By the time we got home, six draft pull requests were waiting for my review. The feature set wasn't due until Monday. It was live on production that afternoon — two days early.


Role
Senior Product Designer
Timeline
Jul – Aug 2026
Team
Design, Product, Engineering
Surface
Policy page

Figures verified as of August 2026

Everyone is trying to make AI delivery fast. Fast is free now. The expensive part — the part nobody hands you a tool for — is making it safe, legible, and adoptable: safe enough to point at a regulated insurance product, legible enough that stakeholders never wonder what the machine did, adoptable enough that engineering stood up a hosting pathway modeled on it and a peer's agent session ran in my repo, under my conventions, unsupervised.

a ~3,500-test suite·seventy-plus pull requests·6 days, mandate to production·~5 hours, stakeholder feedback to production·zero engineering build time — their involvement was access setup and architectural review, by their choice

This is the case study of that system. Not the prompts — the operating system: a public delivery queue, a decision-closure doctrine, and a governance document that supervises AI sessions when I'm not watching. It reduces to three moves, and all three are stealable: derive the truth. Govern the risk. Design for the inheritor.

One thing this story is not: it is not about a designer who learned to code, and it is not about robots doing my job. My first commit wasn't code. It was a conventions document — and that decision, made before a single line existed, is why everything after it worked.

Phase 1 — Earning the standing

The system didn't start with the mandate. It started three weeks earlier, with work so unglamorous it barely looked like anything.

I set one standard for our two-person design team: options ready before anyone asked. Task variations and mobile mocks delivered decision-ready within days, tradeoffs already framed. I took ownership of the service-discovery research, and made one small structural call on a new underwriting tool: information architecture first, edit-screen design deferred until engineering confirmed the technical constraints — the first time design sequenced the work instead of receiving it.

None of it looked like transformation. It looked like reliability. Then the VP challenged design to stand up a working application without engineering, and it turned out reliability was the collateral. The mandate arrived three weeks after the standard did.

Phase 2 — Installing the operating system

Standing bought the right to change the process. The last week of July, I installed two pieces of infrastructure — both documents, both boring on purpose.

The decision log. The new underwriting tool got a published log: nine dated decisions, each with its options, its resolution, and its consequences — plus a section for the risks we couldn't resolve yet, labeled "pending technical verification" so they'd bite as line items instead of surprises. The rule underneath it: every contested question gets written down, framed, and resolved once, on the record. Decisions stopped being re-litigable, because relitigation now had to argue with a dated document instead of a memory.

The same week, that discipline paid for itself in a different room. A tab-structure debate on the policy page looked like a new argument — until I excavated the company's own prior-year redesign research and showed the direction had been paused, not rejected. The project reframed overnight from "new idea needing defense" to "resuming a research-backed direction," with the record doing the defending.

The Line. Design work at the time arrived from three PMs as DMs and drive-bys, every request urgent, "ASAP" functioning as a date. The scarcity was never the problem — the invisibility of the scarcity was. So I built one public, ordered queue for all design work. The house rules are the whole design, and they're stealable as-is:

  • Every row carries two dates: the PM's need-by, and design's target. Both public.
  • New requests enter the line. They don't cut it.
  • To move something up, you name what moves down — and the trade happens in the table, not in DMs.
  • Below the line sits an auto-populated intake, so nothing arrives as a surprise.

I pitched it to the VP over breakfast; he endorsed it the same morning and it went to the PM team that day. "ASAP is not a date" became policy by lunch.

Two moments told me it had taken. The first item through the full lifecycle entered the queue, got a public date, and shipped exactly on it. And about a week in, a PM messaged asking to swap her project's slot with another's — negotiating with the queue, not with me. A private to-do list makes the designer the bottleneck; a public queue makes capacity a shared constraint that PMs trade against each other. After that, the dates started defending themselves.

By August, design had a decision record nobody could reopen and a delivery schedule nobody could jump. That's the whole operating system — and at this point it had nothing to do with AI. Which mattered, because the next week, the VP asked a question that assumed the machine could be trusted with much more.

Phase 3 — The machine goes vertical

In early August, our VP challenged design to stand up a working application without engineering — his framing, roughly: why do you need an engineer for this? It was a fair question. The tools existed. What didn't exist was a reason to trust them.

Because the risky part of AI-assisted delivery was never the code. Agents write competent code. The risky part is everything around the code: what data is allowed to exist in the repo, how changes reach main, what "done" means, and what happens when the machine meets a question it can't answer. Get those wrong and you don't have a prototype — you have confident fiction pointed at a regulated insurance product.

So I created the repository, and the first commit wasn't code. It was a conventions document — an engineering manager in file form. Everything that came after inherited it. The rules, in plain terms, all stealable:

  • All fixture data is fictional. Always. No real customer information under any circumstances — not anonymized, not sampled. Nonexistent. The safest architecture is the one where the dangerous thing isn't guarded; it's absent.
  • The production codebase is read-only reference. Learn from it, never touch it.
  • Every change moves branch → pull request → merge, tests green. No exceptions for speed, because the exceptions are where speed goes to die later.
  • Visual values come from the design file's API — never eyeballed from screenshots. Eyeballing is where drift is born.
  • When an agent hits ambiguity, it escalates a question instead of guessing. This is the rule I'm proudest of, and the one that makes the rest work.

Then the machine went vertical. Working prototype shell inside a week — eight data-driven tabs, alert priority rules, term time-navigation. The project's PM adopted the pipeline as the official workflow, and every handoff disclosed what the agents did and what I decided — because trust in the output starts with legibility about the process. The disclosure wasn't a courtesy. It was load-bearing: it's the difference between a team adopting your pipeline and a team quietly wondering what it should double-check.

Six days after the challenge, there was a production-hosted application — delivered to the VP roughly forty minutes before he walked into an arbitration meeting between two PMs with a genuine prioritization dispute. I'd sent a pre-read the day before: the actual questions under the dispute, a framework for ownership — PMs own the inputs, design owns the design decisions, each PM drives her own lane — and what each resolution would unblock. The meeting resolved in about twenty minutes, on the pre-read's framework, outcomes in writing. Nobody experienced it as losing, which is why it held.

MAT 1The artifact — policy page overview
capture 2272 × ~1500

Policy page overview — the built artifact

Meetings, it turns out, are terrible places for first drafts of decisions and excellent places for last drafts. If the structuring happens before the room, the room's job shrinks to consent.

And the conventions document? I found out it worked about ten days in, when an overnight agent session — unattended — chose the correct branch-and-PR workflow on its own, ran a secret scan before pushing, and, on discovering a mismatch between its instructions and the deployment plan, flagged it instead of silently fixing it. Another day, two concurrent sessions collided in one checkout; the second detected the collision, verified nothing was contaminated, rebuilt its work in an isolated worktree, and documented the incident unprompted.

That's when I understood what the first commit had actually bought. Anyone can make the robots fast. The job is making them safe and legible — and a well-written process document scales in a way attention never will.

Phase 4 — The foreman

By mid-August the pipeline worked, but it had a bottleneck, and the bottleneck was me. One governed agent session per pull request meant every brief waited for me to shepherd the previous one. The machine was fast; the queue into the machine was human-speed.

The Saturday I had six PRs' worth of finalized decisions and a dog who needed walking, I decided the bottleneck was structural, not personal — and structural problems get architecture, not effort.

The foreman pattern is one master document that a coordinating session splits into parallel briefs. Each build agent gets three things:

  • An isolated worktree — its own copy of the codebase, so parallel sessions physically cannot collide.
  • An enforced session name that flows into its branch, its PR title, and its reports — so I always know who's talking, even six conversations deep.
  • The standing rules, with authorities ranked: written session decisions beat mocks, the design file's API beats eyeballing, fixtures own the data — and ambiguity comes back as a numbered ruling request, never a guess.

The rulings mechanism is the heart of it, and it's the part worth stealing even if you never run an agent. The questions that come back are exactly what a good junior engineer would ask — does this render if the reference design doesn't show it? who sees this tab? — and I answer them in batches, in writing. Then the answers become canon that every future session inherits. Most process makes ambiguity recur; this one makes it cheaper over time. The tenth session inherits two hundred settled questions and asks three new ones.

Reports arrive in a standard table — session, branch, PR, test counts, flags — because a foreman who has to read six essays isn't a foreman, she's an editor.

And the merge train closes the loop: PRs land one at a time, rebased, with derived snapshots regenerated from the contract rather than hand-merged — hand-merging a generated file is how subtle corruption survives review — and ancestry verified before anything is called landed. That protocol was written in the blood of exactly one early bad merge. It has not bled since.

That first Saturday, the build phase of six pull requests ran unattended during a 45-minute dog walk. My review and the merges came after, one at a time, on the train. A second stack ran during a golf round that afternoon. The PM's Monday deadline — every tab populated, the full banner system correct — was live on production Saturday, two days early. His follow-up list the next morning closed nine-for-nine the same day, including requests he'd edited into the message after posting, which the intake caught by re-scrubbing the thread — and one request that turned out to contain a genuine pre-existing bug, found and fixed with a regression test.

But the number was never the point. The point is what the job became that Saturday: not typing, not even directing — specifying, ruling, and reviewing. The dog got her walk. The work didn't notice I was gone.

Later, the pattern passed the only test that matters for a process: it worked without me. A peer's agent session ran in the same repository, under the same conventions — chose the right workflow, followed the rules, shipped safely — unsupervised. A process that only its author can run is a talent. One that runs in other people's hands is infrastructure.

Phase 5 — Running it

Sprints make good stories. Systems make good products. What matters most about this one is the part that looks least impressive: how it runs on an ordinary Tuesday.

Truth flows one direction. A working session makes a decision. The decision lands in the mock. The mock's values feed the code — extracted through the design file's API, never eyeballed. Decision → mock → code, a pipeline with a fixed direction, where updating the mock isn't documentation after the fact; it's how a decision propagates. Skipping the mock isn't a shortcut. It's a broken pipe.

That extraction rule has caught things no review would have. Once, a mock lagged behind a decided rename and the brief-writing pass caught it before a line was built — the system detecting a stalled station, which is exactly what it's for. And while ruling a set of status components against machine-extracted design values, a vocabulary discrepancy surfaced: the code carried a status called Superseded that appears nowhere in the design. The design's word for that state is Inactive. Nobody knew "Superseded" was an invention until the design file was treated as a machine-checkable source of truth. That's the difference between a value that was checked and a value that was derived.

Parity is enforced, not maintained. A renamed label in canon fails the build until the code agrees — vocabulary is tested. The state directory doesn't claim completeness; it derives it, walking every fictional policy through every term and role using the actual selection logic. If the spec grows a state the build lacks, the directory tattles. Full-canon coverage was verified in late August: a 23-policy fully fictional cast with production-true shapes, every documented state carrying a living example, and our own audit catching the one gap. Even the demo clock is defended: a test greps the source tree for rogue date calls so the frozen date can't drift. None of this requires anyone to be careful. That's the point. Discipline that depends on vigilance is a countdown; parity here is a property the system asserts.

MAT 2The state directory — drawer open
capture 2272 × ~1500

The state directory, derived rather than declared

MAT 3 · ABanner canon — 4–6 states
capture 1096 × ~700 each

MAT 3 · BBanner canon — 4–6 states
capture 1096 × ~700 each

Banner canon — one state per family, generic copy

The next component is already designed: the same API integration that builds the page can watch it — a scheduled agent diffing the design masters against the built components and filing drift not as fixes but as ruling requests: mock says X, code says Y — which is canon? Because sometimes the code is right and the mock missed its update, and only the human holding canon can say which. Drift is a detection problem, not a discipline problem. That monitor is the designed next step, stated here as designed — this case study keeps shipped and planned in separate sentences on purpose.

MAT 4 · ARole scoping — agent / client
capture 1096 × ~1400 each

MAT 4 · BRole scoping — agent / client
capture 1096 × ~1400 each

Role scoping — agent and client views of one fictional policy

MAT 5 · AFixture cast — index + naming crop
capture 1096 × ~1400 each

MAT 5 · BFixture cast — index + naming crop
capture 1096 × ~1400 each

The fictional cast — every policy, and a naming detail

And a running system can still burst. At a 9am feedback session late in the month, an agency-facing stakeholder pushed hard on one structural change: get the additional-parties information — mortgagees, additional insureds — off the overview and into its own tab. He was right, and he was firm. Historically, "you're right" starts a two-week cycle. The change was on production about five hours later — and the interesting part is what filled those hours. The build ran in two legs, deliberately. Leg one was pure reconnaissance: a read-only pass over the production codebase to learn what party types actually exist, what fields each carries, and what realistic frequencies look like — derived from the code's own structure, never from customer records. The recon killed two pieces of fiction on the spot: a party type our old component implied that doesn't exist in production, and a payment-flag nuance narrower than we'd been rendering. It surfaced four genuine product questions as written rulings — one of which I escalated onward, because it was the PM's to make, not mine. Leg two built it: new tab, the fictional cast recalibrated to the derived ratios with the basis documented in the pull request, two hundred new tests.

The compressed timeline is the flashy part. The durable part is that the request came out more truthful than it went in.

And it moves decisions at a different rate. A comparable project under the old relitigation-heavy process spent roughly a quarter in design churn before its interface reached QA. The policy page absorbed a comparable density of contested decisions in roughly a month. Same class of work, same organization, overlapping stakeholders. The variable that changed was the operating system.

The system travels. The same month, the same operating system ran a second surface: an underwriting rule editor — different product, different PM, different stakeholders — carried a full lifecycle from draft through active status and demoed successfully to underwriting, with the next scope agreed in the room.

Then it did something better than travel: it got rebuilt from scratch in a day. A second delivery repository went from nothing to complete infrastructure in one working day — production hosting with a deploy gate running before every build, a component workshop with visual-regression baselines on every pull request, and the first project's process rules ported wholesale so the new repo started with every lesson the first one had learned by incident.

The proof it worked arrived immediately. The very first deploy failure was the gate catching nineteen test failures that continuous integration had passed. The cause: the UI framework's new major version ships a testing utility in development builds and strips it from production builds, so tests that passed in development-flavored CI failed against the real production build. Root-caused in a fresh clone, the dependency pinned, CI upgraded to test both configurations. And the same latent exposure was written up as a finding for the production codebase — the new system caught a bug the older, larger one had been quietly carrying.

A pattern that runs once is an anecdote. Run twice, on surfaces that share nothing but the method, it's a system. Rebuilt from its own documentation in an afternoon, and catching a defect in the codebase that predates it, it's infrastructure. And it had already passed the harder test: a peer's agent session, running in my repository under my conventions, unsupervised, safely. Infrastructure is the process that works in other people's hands.

Terminal output of megancasebier.com's own publish gate refusing a staged breakage
The gate refusing — this site’s own publish gate, mid-refusal, on a staged breakage

Part Two is the handoff

Here's where the story stops being finished.

The prototype is the plan of record for the production rebuild of the policy page, and an engineering pod is forming around it. The question that phase will answer is the one this whole system was built to earn: what happens when design's code meets the production codebase?

Part of that answer is already agreed, in writing. On a second surface — an underwriting workbench — design authors the display layer in the production team's own file convention: design owns the pure display components, the pod owns the containers and the data wiring. I merge freely in my own repository behind my own gates; the formal review happens on their side, when a developer copies each chunk into the production repo and opens team code review, with chunks sized to stay under their five-hundred-line review limit. Sequencing is leaf-first — smallest components first, each validated as it goes — with strict oversight that relaxes as trust is earned. Agreed in thirty-five minutes.

Where it actually stands, stated precisely: the first chunk — a set of status components — was built, merged in the design repository, deployed, and handed over with an extraction manifest listing exact files, dependencies, and a pinned source revision. First component delivered for adoption; production merge pending. The agreement is executed to the point of first handoff, not yet to production merge.

The convergence work is underway. The data layer already speaks the production API's shapes — it has since the adapter's first byte-identical round trip. The component layer is being mapped to the production component library now, one governed workstream at a time. The repository stays alive after the handoff as the living design surface, with the drift monitor watching parity between what design decided and what production renders.

Design didn't ask for a seat at the table. It built a delivery system, shipped a product with it, and made the code worth inheriting.

Part Two is what happens when engineering inherits it. It's being written now.