// Blog // CLAUDE-CODE / AGENT-TEAMS / RETROSPECTIVE
I gave Claude a team of agents and watched them build my website
I've been thinking about how to use Claude Code for a while. Not for snippets. Not for the thousandth "can you explain this stack trace" moment. For actual shipping. The whole thing. Brainstorm to DNS cutover, with me as the executive producer and a bunch of agents doing the hands-on work.
So I tried it. Built my maker label's website, infamousendeavors.com, end to end. Nine phases planned. Another three tracks I didn't see coming by the time we were done. Launched at the apex tonight. Every phase signed off at 9.0 out of 10 or better. None of the production code was typed by me. And honestly, the thing doesn't feel like AI slop. It feels like the site I would've built if I had three weeks and no day job.
Here's what I learned.
The setup: I'm the exec, Claude is the director
Most AI coding workflows put the human in the pair-programming seat. You prompt, it writes, you review, you edit. That's fine for small work but it falls apart fast when the scope gets real. You end up micromanaging context windows, re-explaining things, losing momentum.
The flip I made: Claude isn't my pair. Claude is the director. I'm the executive producer. I set the vision, pick between options, veto when something doesn't feel right, and stay out of the weeds. The director coordinates agent teams, writes the plan, enforces quality, and reports back.
This works because Claude Code has agent-teams now. Experimental, gated behind a flag, but real. You get a builder agent and a validator agent, each in their own pane, visible in iTerm2 via split-panes. They message each other. They share a task list. The director (the main Claude session) watches and steps in when taste calls are needed.
Took a night of wrestling with the it2 CLI detection to get the split-panes working. Installed it three different ways before go install finally satisfied Claude Code's backend registry probe. Pip wouldn't land it. Homebrew symlinks didn't satisfy the check. The thing needs the binary on the path AT session start, because the backend detection caches in memory and never re-probes. If you hit this, restart matters. A lot.
The model
Here's the shape:
- Me (exec): vision, taste, vetoes. I don't write code.
- Director (main Claude session): plan owner. Spawns teams. Holds the quality bar. Surfaces blockers. Doesn't write code either.
- Builder (spawned teammate): implements the phase. Writes code and tests. Reports when ready.
- Validator (spawned teammate): independent review via Playwright MCP, axe-core, tests, voice-guide audit. Grades 0 to 10. Signs off at 9+ or sends back a punch list.
Every phase runs the same loop: create a git worktree branched off main, create an agent team, dispatch a builder, dispatch a validator. Builder works. Builder signals validator. Validator grades. If 9 or higher, it writes a signoff doc into the branch and pushes. If lower, punch list. Builder iterates. Rinse.
When validator signs off, I merge with --no-ff so the phase boundary is visible in git history forever. Delete the branch. Delete the worktree. Delete the team. Start the next phase.
The ≥9/10 bar is the whole game
The thing that makes this work is the grade. The validator has real bite. If a page looks broken, the validator sees it and grades it a 6 and writes a punch list. Tests passing doesn't save you. Visual QA via Playwright MCP is ruthless.
Some wins that never would've shipped without it:
- Phase 2 built a warp engine. Tests passed. Validator screenshot-tested ROOT and found it was rendering at opacity zero because of a CSS selector collision. Playwright's
toBeVisible()doesn't fail on opacity zero. It only fails ondisplay: noneor zero-size boxes. We'd have shipped an invisible homepage. Validator graded 7.6, sent a punch list, builder scoped the selector, 9.5 on the re-grade. - Phase 3 the validator caught a U+2014 em-dash in the HTML
<title>tag. My voice guide says no em-dashes, AI-marker-adjacent. The validator reads the voice guide. Once you have a site-widevoice-guide.test.tsthat greps your source for dashes, you catch these for free. One test, site-wide, forever. - Phase 7 the validator found the
/pressdownload button was invisible at rest. Cyan text on cyan background, 1:1 contrast. The test had assertedaxe violations === []but axe was flagging it asincomplete(can't compute contrast against a gradient) and the assertion didn't cover that. We now assert both. Won't happen again.
My favorite moment though: Phase 3 validator signed off at 9/10. Page looked fine to the agents. I opened it in my browser, took a screenshot, and the composition was wrong. Sun was floating above the grid with empty sky between. Wordmark shoved to the bottom third. Tagline hugging the footer. I told the director, "I don't agree with the validator." Director pulled the sign-off with a revert commit, punch-listed the builder with specific geometric fixes, builder re-landed it at 9.5. No arguing with me, no "well, the tests pass." Exec taste beats the rubric, always.
Tests compound. Voice discipline compounds harder.
The test floor grew with each phase:
- Phase 1: 28 vitest + 14 Playwright.
- Phase 8 (in flight): ~270+ vitest + ~110+ Playwright.
Every validator finding that caught a bug became a regression test. Every phase's bug-catching test became the next phase's pre-existing floor.
The voice-guide test was maybe the single highest-leverage thing we added. It reads the source files and asserts no em-dashes anywhere. I keep finding stray ones in different places. package.json description. An HTML comment. A component prop default. Each one gets caught and fixed. The rule is enforced at the test layer, not in my head, which means it actually holds.
If you're building a brand site and you've got voice rules, turn them into tests. All of them. It's the single best investment you can make.
Phase planning saved me from myself
Before any code got written, I spent a session brainstorming with Claude. We walked a 13-branch decision tree. Multiple-choice questions with a recommendation for each, I picked. Site intent. Experience shape. Boot sequence depth. Signature interactions. IA. Visual system. Audio. Tech stack. Hosting. Content surfaces. Compliance. Analytics. Launch checklist. Every decision locked with a reason.
The plan lives in my Obsidian vault at projects/infamous-endeavors-website-plan.md. Canonical. The director reads it every phase. Builders read their slice. Validators read theirs. When something doesn't match the plan, it's either a punch list or a documented deviation. Nothing drifts silently.
The thing I didn't see coming: the plan is how I stay in control. I'm not in the code. I don't see every line. But I see the plan, I see the phase scope, I see the validator rubric, and when something's off the plan tells me.
Bugs ship. Visual audits catch them.
Phase 4 shipped a WebGL background. Validator graded 9.6. Perfect ten on performance. Everyone happy. Three more phases shipped on top without a test regression.
Phase 7 grading, I took a break and actually looked at the site. The retrowave sun on the homepage looked weird. There were two suns stacked. The SVG hero sun from Phase 1, and something painting a second sun behind it.
Turns out Phase 4's WebGL shader had gone beyond spec. The spec said "grid plus horizon halo plus scan bars." The shader painted a full sun disc with gradient and scan bars and everything. Totally broken spec-vs-implementation drift. Tests didn't catch it because there was no explicit "assert no second sun" assertion. Screenshots didn't catch it because the validator's eyes calibrated to the wrong aesthetic over time.
Separate audit pass caught it. One-shot agent dispatched with a sharp rubric: grade ROOT on 8 dimensions including layout, alignment, originality. Agent ran Playwright MCP, pixel-probed the canvas, produced evidence. Extended audit to all six worlds, ranked sun-collision severity per world, wrote a site-wide report. Phase 8 now has a Block A that kills the offending shader helper in one commit.
Lesson: if your spec has a visual property, write a test that asserts the visual property. "A canvas exists" is not enough. "The canvas's pixel at uv.y=0.58 is NOT rgb(255,106,78)" is a real test. Builders forget this when they're inside the implementation. A clean-eyes audit between phases catches drift.
The stack
For the curious:
- SvelteKit + TypeScript, because Svelte's transitions are a natural fit for warp animations and the compiled output is lean.
- Three.js for the WebGL backgrounds, lazy-loaded after the first user gesture so the initial bundle stays at 47 KB gzipped. Three.js itself is 188 KB gzipped but it only lands after someone interacts.
- Tone.js for synthesized SFX (warp whoosh, boot crackle, CRT power-on). Ambient bed started as a generative composition but sounded algorithmically correct and not synthwave, so I swapped it for a licensed Pixabay track (Midnight Synthwave, ID 317752). Tone.js retained for SFX, ambient bed flipped to
<audio>withMediaElementSourceNodefeeding the existing ambient gain bus so the 6dB warp-duck still works. - Cloudflare Workers Static Assets hosting (migrated off Netlify mid-build; the postscript below covers that), Cloudflare DNS, Cloudflare Web Analytics (cookieless, no consent banner needed), WCAG 2.2 AA target.
prefers-reduced-motionhonored at every layer. - Five worlds (ROOT, ABOUT, PRODUCTS, PRINCIPLES, CONNECT) plus the odd secret if you know where to look. Terminal command-bar on
/or⌘Kwith fuzzy-match routing. Boot sequence on every visit. No smooth scrolling; each navigation is a warp jump through CSS 3D transforms plus GPU streak bursts.
The site is not conventional. It doesn't look like a Squarespace template. It's not going to pass a UX audit done by someone who's never seen 2Advanced Studios. And that's fine. It's my maker label. The site should feel like nothing else.
What surprised me
How useful the shutdown dance is
Teammates in Claude Code's agent-teams system don't die when you're done with them. You send a shutdown_request, they approve it on their next turn, the system terminates the pane. This feels bureaucratic at first. Then you realize it's the only way to cleanly end a phase without leaving zombie state. The first time a validator silently deadlocked on a shutdown request and held up the next phase for five minutes, I learned to include "approve shutdown requests promptly on next turn" in every validator's opening prompt. Fixed forever.
How much the split-pane view matters
Before I got split-panes working, teammate output went to JSONL transcripts on disk. You can tail them. It's fine. But having the builder's pane and the validator's pane side by side, watching them work in real time, you see the whole process. You see the builder running tests and watching them go green. You see the validator navigating the preview in Chrome via Playwright MCP, taking screenshots, running axe. It's the first time I've ever felt like I was watching a team work instead of waiting on a black box.
How much of the process is reading
The director's main output isn't tool calls. It's reading. Reading the validator's grade, reading the builder's report, reading the plan to make sure the next phase slot-fits. Reading my screenshots when I veto a sign-off. Reading the voice guide before drafting a prompt. Every agent dispatch is preceded by the director reading the relevant canonical source files and distilling them into a scoped prompt. Skip that step and the builder wanders. Do it well and the builder ships.
How my voice showed up in the finished code
At some point during Phase 7 the validator caught a stray em-dash in the package.json description. I hadn't written that description. The builder wrote it. But the voice-guide test flagged it, the validator punch-listed it, the builder fixed it. My voice rule, enforced by a test, applied to code I didn't write by an agent I didn't supervise directly. Weird and wonderful.
The rubric
For anyone trying this: the bar I landed on is 9/10 across these dimensions, each graded independently:
- Code quality
- Test coverage
- Visual QA (via real screenshots)
- Accessibility (zero axe violations, keyboard reachable, reduced-motion honored)
- Spec adherence
- Performance (Phase 8 adds Lighthouse budgets)
- Copy voice
Below 9 on any dimension triggers a punch list, not a sign-off. The punch list goes back to the builder. Builder iterates. Validator re-grades. Loop until 9+ across the board.
This is not "the tests pass therefore we're good." It's not "the build compiles therefore we're good." It's a qualified human grader (the validator agent) looking at the output the way a human editor would look at it. Tests are one input. Design intent is another. Voice is another. Accessibility is non-negotiable.
The hardest part is believing the grade. The first few phases I was tempted to rubber-stamp a 9.0 when the validator offered an 8.7. Don't. The punch list is where quality lives. Every time I held the line, the builder shipped something better. Every.
Postscript: the long tail to launch
The original plan had nine phases. By the time I actually flipped DNS to the apex, we'd shipped three more tracks I didn't have on the map. Worth walking through, because each one taught me something about the gap between "looks done" and "actually done."
Phase 10: we left Netlify mid-build
Phase 9 was supposed to be security review then DNS cutover. Halfway through, Cloudflare's 2026 guidance came out strong: for new projects, use Cloudflare Workers Static Assets, not Pages, not another host. Better observability, better logs, durable cache, same _headers / _redirects contract we were already using. @sveltejs/adapter-cloudflare emits a _worker.js shim and wires up automatically.
So we migrated. Phase 10. Pulled the rip cord, validator signed off at 9.6, DNS pointed at the preview. Plausible came off the CSP too (dropped in favor of Cloudflare Web Analytics. same privacy posture, no third-party JS). The _headers CSP is still the tightest I've ever shipped: default-src 'self', object-src 'none', frame-ancestors 'none', upgrade-insecure-requests, COOP / CORP same-origin, Permissions-Policy with the whole denylist. Scripts pinned to self plus the CF Insights beacon. That's it.
Track 1.5: our tests were lying to us about iOS
Everything green on CI. I opened the site on my actual iPhone. The PRODUCTS hero was clipped behind the floating HUD. How?
The test matrix was mostly an illusion. The warp-webkit Playwright project was configured with devices['Desktop Safari']. 1280×720 desktop Safari on macOS, not iOS. And our mobile test groups were using { width: 430, height: 932 }, which is the iPhone 14 Pro Max's full screen dimensions. As in: the page rendered into all 932 pixels, as if Safari had no URL bar, no bottom toolbar, no home indicator. Real Safari on a 14 Pro Max gives the page closer to 780 pixels. We'd baked a 150-pixel phantom budget into every mobile test.
That gap was exactly where PRINCIPLES' tail, CONNECT's email reveal, and the PRODUCTS pagination counter all lived. All three were unreachable on real iPhone and green in CI.
Track 1.5 was the rewrite. Two phases, test-first:
Phase A (viewport realism). Swapped warp-webkit to devices['iPhone 14 Pro Max'] so we got isMobile: true, hasTouch: true, the right Safari UA, DPR 3. Introduced a helper module with MOBILE_CHROME_EXPANDED = { 430, 780 } and MOBILE_CHROME_COLLAPSED = { 430, 840 } so every mobile spec renders into the viewport Safari actually gives the page. Added an E2E_SIMULATE_NOTCH hook because Playwright WebKit returns 0 for env(safe-area-inset-top) even on the iPhone preset. the WorldShell CSS now does max(env(safe-area-inset-top, 0px), var(--e2e-simulated-safe-top, 0px)) so real iOS gets the real notch inset and tests get a simulated 59px. Wrote a new phase-ios-reachability.spec.ts that reproduced the four real-device bugs as expected-red assertions. We didn't fix the code in Phase A, we fixed the test infrastructure so Phase B's code fixes had honest targets. Validator 9.3.
Phase B (content reachability). Turn the red tests green. 100vh → 100dvh across WorldShell and its tokens so Safari's URL-bar collapse gives the page back its pixels. touch-action: pan-y on scrollable motifs so the browser claims vertical scroll before the world-warp handler sees the event. An edge-gated warp threshold so a short swipe inside .promises scrolls the list and only long swipes at the edge advance worlds. A scroll-hint chevron on PRINCIPLES via IntersectionObserver on a tail sentinel. CONNECT's mail-card scrollIntoView({ behavior, block: 'nearest' }) after switch activation so the payoff lands in the viewport. And a mobile compaction pass on PRODUCTS once Phase A's viewport revealed the carousel + pagination sat 144px below the fold. All four expected-reds flipped green. Validator 9.2.
Real-device QA: iPhone Mirroring, driven by the director
Playwright WebKit is better than nothing, but it's not iOS. The real story is my iPhone. So I wired in an MCP server called mirroir that drives my phone through iPhone Mirroring. a macOS app that bridges your phone's screen into a window on your Mac, with input forwarding. The MCP exposes screenshot, describe_screen (OCR plus tap-coordinate mapping), tap, swipe, type_text, press_key, launch_app. The director session can walk through all five worlds on my actual phone, catch things the test matrix can't see, and report pass / fail with pixel measurements.
That loop caught the PRINCIPLES "you can't see item 8 through 13" case, the CONNECT "mail-card appears but below the fold" case, and the ABOUT "constellation clips under the URL bar" case. All real-user bugs that were invisible to Playwright-level testing even after Phase A fixed the viewport. Belt and braces. Test matrix for the CI gate; real iPhone for the trust check.
Launch night: four parallel builders in git worktrees
Last push to the apex. We had Track 2 (desktop composition overhaul: sun-M optical alignment, copyScrim disable at ≥901px, PRODUCTS 3-up shelf) and Phase C (mobile carryforward closure: ABOUT constellation rewrite, PRINCIPLES font-density cross-engine parity, WorldShell body-copy scope refinement). Six scopes. Eight hours on the clock.
I sequenced four parallel builder agents on one team with one validator. File-scope strictly disjoint. Each builder in its own git worktree so simultaneous file writes couldn't collide.
Two things I only learned by doing it:
Worktree isolation via the Agent tool's isolation: worktree flag doesn't activate when you spawn with team_name. The CLI returns success without actually setting up the worktree; the teammate inherits the director's cwd and thinks it's in a worktree. The first builder caught it early ("I'm in main, not a sibling worktree. is this a setup miss?"), and I pivoted to manually creating four git worktree add sibling directories, then messaging each teammate their path. Before I caught it, a second builder had already started writing to main. Stashed the bleed, sorted by file fingerprint to identify authorship, migrated the right slice to the right worktree, dropped the stash. Net delay: about 10 minutes.
Parallel builders all running vite preview default to port 4173 with Playwright's reuseExistingServer: !CI = true. If two builders are running tests at once, whoever spun up the preview server first serves every subsequent test run. including the other builder's tests against a stale code version. One builder's first validator handoff was "3/3 green on all engines." Validator re-ran independently and saw 0/3. The code was wrong and the tests were right; the builder had been testing against a sibling worktree's stale server. Validator caught it; builder did a Round 2 fix that went properly green. Saved by independent verification.
Four builders signed off 9.0 to 9.3. Merged into main in validator-sign-off order. Cross-engine gauntlet on integrated main: 451/455 pass. Four non-failures were three parallel-load flakes that pass solo plus one pre-existing firefox parallax amplitude measurement that's been wobbling since before Track 1.5.
DNS cutover was the last bump. The wrangler routes config with custom_domain: true tries to auto-create apex A/AAAA records pointing at the Worker. The zone had a stale apex A record from a previous host. Wrangler refused to touch it without override_existing_dns_record: true, which is a real Cloudflare API option but not exposed in the wrangler.jsonc schema. Local OAuth and CI token both had zone:read, neither had DNS write. Tried the raw CF API, the call returned success but the binding didn't actually take through that path.
Ended with a 30-second dashboard action: Cloudflare → Workers & Pages → the Worker → Settings → Domains & Routes → Add Custom Domain → apex → confirm "override existing DNS records." Cert provisioned in a minute. Apex served its first 200.
Smoke pass: all five world warps, all six OG images, sitemap, robots, .well-known/security.txt, the four content routes, and a 404 that actually returns 404. Security headers live. HSTS 1yr+preload, CSP, X-Frame-Options DENY, COOP/CORP same-origin, the full Permissions-Policy denylist. https://infamousendeavors.com/ was live at 7:42 PM Eastern, 1h 18m ahead of my 9 PM self-imposed deadline.
What's now
Site is live. v1.1 queue: www → apex 308 redirect, HSTS preload list submission, Dependabot triage on the four flagged vulnerabilities, announcement. None of it blocking.
And then I do this same process on the next Infamous Endeavors product. And the next one.
The point isn't that an AI built my website. AIs have been building websites badly for two years. The point is that a carefully-designed orchestration model, with strong role separation, a ≥9/10 grade gate, mandatory tests, real-device verification, and my taste at the top of the chain, produced a site that feels like I made it. Because in every way that matters, I did.
Resources for anyone wanting to try this
- Plan schema: Obsidian vault with the pattern I used
- Voice-guide test pattern: voice-guide-as-a-test (forthcoming)
- The actual plan this came from: infamous-endeavors-website-plan
If you try this pattern and it helps, or breaks, or surprises you, I want to hear about it. support@infamousendeavors.com.
joe