Skip to content
shuffl
← All articles

· Shuffl team

243 agent threads in a week: how I rebuilt Shuffl from my phone

One person, no laptop: 243 agent threads, 556 merged pull requests and 45 computer-use jobs. The setup behind it, and how to copy it for your own team.

Short answer

Since September 21, 243 Claude Projects threads have built and shipped Shuffl while I dictated from my phone, with up to 127 commits a day. The agents run CI in their own containers, merge and release. When a job needs a login, an agent on my Mac does it with computer use. Here is the setup and how to copy it.

What does one person with agents ship in a week?

I rebuilt Shuffl by talking to coding agents on my phone: on a sailboat, lying in bed, dictating on the drive to daycare. I never opened my laptop. Here is what that added up to by September 28.

agent threads since September 21
243
threads started in 38 seconds
10
merged pull requests
556
commits on the busiest day
127
computer-use jobs on my Mac
45
pull requests a person reviewed
0

Commits to main per day

Commits to main per day, September 9 to 27, 2026, Pacific time. Before September 22 the most in one day was 63. From September 22, when I moved to Claude Projects, there were 63, then 127, 45, 72, 45 and 25.

Commits to main per day in September 2026, Pacific time, from git log. Each commit is one squashed pull request.

What does a prompt look like?

All day I rattle off thoughts into voice-to-text, one after another. Build this. Go check the page that just went live. That button is wrong, here’s a screenshot. I don’t wait for anything to finish. A coordinator agent decides whether each message is new work or belongs to a thread that’s already running, and routes it there. It’s an ADHD brain’s dream.

The prompts are rough. This post started as the message below, “use in physical” and all (it meant Infisical). From that one message the thread found its evidence in the repository and GitHub, wrote the post, drew the diagrams, screenshotted the page on desktop and phone and opened the pull request.

The message that started this post, dictated
Maybe we actually write another blog post next to this that kind of dives deep into how we did this um, technically. So we can like show this off from an engineering perspective. […] make it technical add like diagrams um, you know talk about how Astra used computer use to create new accounts and create API keys and use in physical […] I think that’s like the huge takeaway.

Where each thought goes

I dictate a thought into Claude Projects. A coordinator agent reads it. If it is a new idea, it starts a new thread, which builds, tests, screenshots and merges the change, and it goes out with the next release. If it is about work already running, the coordinator forwards it to that thread, which changes course mid-task and updates the same pull request.

  • Needs me
  • An agent does it
  • Done
I never wait for a thread to finish before sending the next thought.

How do 243 threads run at once?

Every thread gets a fresh cloud container. A start-up hook installs Node, starts PostgreSQL 17 with pgvector and seed data, and gets Chromium ready, so the agent can run the real app and the whole test suite. It builds the change, screenshots every screen it touched, pushes, drives the pull request green, merges it and sends me the screenshots.

Decisions go into shared project memory. When I say “every new feature ships with analytics events” once, every later thread already knows.

The whole setup

I talk to Claude Projects from my phone. The coordinator starts a thread for each message, and each thread runs in its own container with a database and a browser. Ten threads once started within 38 seconds. Threads push pull requests whose checks already passed; main runs the full suite once an hour and releases to production. When a thread needs a login, it files an issue labelled computer-use, and an agent on my Mac works through Stripe, Supabase, Vercel, Slack, Infisical, npm or Google in my signed-in browser.

The five titles are real threads from that 38-second burst on September 22.

Does it stop when I sleep?

Long jobs keep going after I stop talking. On September 23, 57 commits landed on main between midnight and 6 AM.

When the commits landed

Commits to main by hour, Pacific time, Monday September 21 to Sunday September 27. The busiest block is Wednesday September 23 from midnight to 6 AM, with 57 commits, and there are evening peaks most days.

Commits to main by hour, Pacific time, from git log. The agents merge their own pull requests, so a commit lands whenever its checks pass.

What came before Projects?

The first two weeks ran from one Apple Notes checklist, and I never logged into GitHub. One checkbox just said “Celebrations”. Codex turned it into an issue with acceptance criteria and a full implementation prompt, and nine days later the feature was in the product. By September 25 the note had 229 checkboxes and 144 were ticked. The catch was that an idea could wait a day for the next sync, and Projects removed that wait.

Before Projects: Apple Notes as the backlog

I dictate a line into the Shuffl 3.0 note. Every day at 5 PM Pacific a scheduled Codex task reads the note, creates a GitHub issue for each new checkbox and writes the issue link beside it. A worker picks up the issue and merges a pull request. The next sync checks the merge and its evidence and ticks the box in Notes.

  • Needs me
  • An agent does it
GitHub stayed the record. A ticked box in Notes never counted as done on its own.

How do you make a repo agents can work in?

Every thread starts cold, so whatever it needs has to be written where it will look or enforced by a check it can’t skip.

  • An AGENTS.md with the working agreements: product rules, UI rules, where each kind of test goes, how to run the app.
  • One Next.js app with the business logic in one shared package, so there is one obvious place for everything.
  • Append-only migrations, 148 so far. Each agent renumbers its migration to the next free number right before it merges.
  • Rules written as tests. A test fails when an API mutation has no analytics event, and the push is refused when a screen changed but its docs screenshots didn’t.

How fast is the CI loop?

With nobody reviewing code, the checks are the review. At first every push ran them on GitHub’s runners, about 1,000 billed minutes a day. So a thread moved CI into the agents’ containers, which already had the database and toolchain. The pre-push hook now runs everything in about seven minutes, and main releases itself when its hourly checks pass. When main goes red, the coordinator starts a thread to fix it.

Where each check runs

Before a push, the hook in the agent’s container runs the SDK contracts, typecheck, format and lint, a secret scan, both migration rehearsals and the unit and service tests, in about seven minutes. On GitHub the pull request runs only a one-minute gate. After a merge, main’s checks start at most once an hour: one production build, then the HTTP contracts and critical browser journeys against that build, then an automatic release when everything passes.

  • An agent does it
  • Done
Dark mode and the full browser suite run every night, outside this loop.

Why is computer use the biggest lesson?

After the first week, code stopped being the slow part. Everything behind a login was: new accounts, API keys, console settings. Cloud agents can’t sign in as me, and I don’t want production secrets in their containers.

So when a thread hits a step that needs my accounts, it files a GitHub issue labelled computer-use that another agent can follow without me. I point Astra at the label and it works through the queue in my signed-in browser. One thread even kept the Mac awake for a week so the jobs could run overnight. Some of what the 45 jobs did:

A computer-use issue, the shape I reuse
Goal: one sentence on what should be true when you are done.

Rules: use my signed-in browser. Stop and ask at any login, 2FA or passkey prompt. Keys go from the provider’s copy button straight into the secrets manager, never into a screenshot, chat, file or GitHub.

Steps: numbered, one click path each.

Verify: how to prove it worked without printing a secret.

Report: comment here with what changed and no secret values.

How a computer-use job runs

A thread reaches a step that needs my accounts. It files a GitHub issue labelled computer-use with the goal, the rules and numbered steps. On my Mac, Astra reads the issue and works in my signed-in browser. It stops and asks at any sign-in or 2FA prompt. It copies each key from the provider straight into Infisical, updates Vercel and redeploys, then checks the result. It reports back on the issue without any secret values, and the thread carries on.

  • An agent does it
Anything that can’t be undone, like a live payment or a publish, waits for me to type yes.
  • Set up Stripe’s products, prices and customer portal in test and live mode.
  • Made a read-only Stripe key and a read-only database role, so cloud threads can look at production without changing it.
  • Created the npm organisation, published the SDK and CLI, and turned on trusted publishing.
  • Put keys into Infisical, checked they reached Vercel and redeployed.
  • Exported all 35 tables of the old Shuffl database and checked each one against hashes and row counts.
Code stopped being the slow part. Everything behind a login was, until computer use.

What is the biggest computer-use job?

The biggest job is moving the old Shuffl, its own servers and database, onto the new platform. Cloud threads wrote the exporter, the converter and an import runner that can verify, go live and roll back, and the steps that need my accounts run on my Mac.

After the export, a read-only comparison against the live database found timestamps shifted by a time zone. The agent marked that export archive-only and asked for a corrected one. I wouldn’t have caught that by hand.

How do I copy this setup?

If I started over, I’d set it up in this order:

  1. Write an AGENTS.md for an engineer who has never seen the repo, then turn its most expensive rules into tests.
  2. Add a start-up hook that gives every agent session your database, seed data and a browser.
  3. Run CI in a pre-push hook inside that session, and keep GitHub to a fast gate plus an hourly check on main.
  4. Send every thought as its own message and let a coordinator route it. Let the agents merge once the checks pass.
  5. Turn every job behind a login into an issue shaped like the template above, and let a computer-use agent work the queue.
  6. Keep a short list of things that can’t be undone, and make each one wait for a typed yes.

Frequently asked questions

Which agents rebuilt Shuffl?

Codex and Claude wrote the code, with the Fable and Opus models doing most of the Claude work. Astra ran the computer-use jobs on my Mac.

Does anyone review the code?

No person reviews it. Every change passes the checks in the agent’s own container before it can be pushed, then main’s full suite before it can reach production.

Is it safe to let an agent use your browser?

It works under written rules. It stops at every sign-in, 2FA and passkey prompt, keys go straight from the provider into the secrets manager, and anything that can’t be undone waits for me to type yes.

Can I copy this setup?

Most of it is files in a repository: an AGENTS.md, a start-up hook that brings up your database and browser, a pre-push hook that runs your CI, and an issue template for computer-use jobs. The last section lists the order I’d set them up in.