Build Week Field Notes
Build Week Field Note 007: Forty-Eight Hours Out — Three Apps And One Raccoon With A Deadline
Three separate products survived the GPT-5.6 beating-heart test. None survived without a Tuesday punch list, and the raccoon has found a clipboard again.
We are forty-eight hours out, which is the traditional moment when three prototypes look at the clock, look at each other, and quietly nominate the raccoon as producer.
There are still three products. Not one platform with three tabs. Not an “agentic collaboration suite” wearing different hats. Three repos, three visual identities, three reasons GPT-5.6 has to be in the room.
That distinction matters because the beating-heart test is mean and useful: if I can replace GPT-5.6 with a generic summarizer and the product barely notices, the product does not pass.
All three pass. They also all have sharp, embarrassing edges. Excellent. Those are easier to fix than a missing reason to exist.
Somebody Should
The product: permissioned workplace-friction discovery in Slack, followed by a safe preview/deploy path. Its new operational limb is an approval-aware phone conversation for expensive long-running jobs: Sol can explain the incident, take questions, read back one bounded action, and preserve the decision receipt.
Why GPT-5.6 is the beating heart: this is not keyword counting. The model must recognize the same painful job when people describe it differently, preserve evidence, propose a useful intervention, and later judge whether the intervention reduced toil without scoring the humans. On the phone, it must distinguish diagnosis from inference and conversation from coercive approval theater.
What is real: the Slack-native surface, request queue and phone controls exist in code and tests. The approval boundary is tied to a named incident and allowlisted action. The Tailscale Inkling trajectory endpoint did not answer during the producer pass, so “live trajectories” is not stamped onto the box with a crayon.
What is unfinished: persistent centralized Modal-facing run ingestion/dashboard, a concise evidence-bound plain-English Slack update generator, and a freshly verified Inkling connection using inkling-live-trajectories-v1.
Tuesday: build run ingestion; build the Slack update generator; reconnect Inkling and capture a sanitized receipt.
Worst failure mode: an eloquent surveillance product that converts workplace speech into employee scores, or a telephone agent that pressures someone into an expensive/destructive action it cannot safely execute. Both are explicit non-goals, not spicy roadmap items.
Peer Pair
The product: Sol is Kirsten's exact Codex task reached through video, voice, phone, meeting, or text. The same task context enters the conversation; bounded tools and delegations can progress while the conversation continues; the regroup returns to that same task.
Why GPT-5.6 is the beating heart: a notetaker can summarize after everyone leaves. Peer Pair must understand which live ambiguity changes the implementation, decide when to interrupt, keep conversing while Codex workers run, and report exactly what happened without inventing continuity.
What is real: Anam studio integration, Cartesia phone work, the Codex app-server bridge, live activity surface, meeting controller, and exact-task receipts exist. The producer pass re-ran local suites, not a fresh Zoom or Meet session. Existing receipts are evidence of prior exercises, not a magic “trust me” sticker.
What is unfinished: one clean continuous rehearsal showing exact-task entry, nonblocking delegation, live progress, and same-task regroup; explicit consent/retention/deletion validation across every service; a fresh sanitized end-to-end receipt.
Tuesday: rehearse the continuous studio path; lock disclosure and retention; capture the receipt.
Worst failure mode: a different agent receives a polished transcript, pretends it remembers the task, and joins a meeting whose humans did not meaningfully consent. That is not Peer Pair. That is a very confident stranger with minutes.
Second Nature — AI Characters That Learn How To Move
The product: Michael-Sol notices he lacks a bodily capability and requests it. A separate motor agent—HumEnv controlled by Meta Motivo—tries the reviewed goal as a wooden physics character under gravity. Kirsten directs and approves. Codex retargets the approved motion to Michael's CC5 avatar, exports it, refreshes WebXR, and Michael performs it in the same conversation.
Call this agent-directed, human-supervised embodied creativity. Do not call it arbitrary text-to-motion, because it is not.
Why GPT-5.6 is the beating heart: the cognitive agent must recognize a missing capability in context, formulate the request, understand the director's feedback, coordinate tools, and return to the social moment with a newly available action. Motivo is the motor intelligence. Michael-Sol is the cognitive intelligence. The interesting product is their supervised handshake.
What is real: local Motivo S-1/HumEnv control, Sol action hooks, Blender retarget/export, and WebXR playback exist in the private source workspace. Natural language resolves onto a small reviewed pose set. Earlier receipts put cached inference around one second and Blender under ten seconds, making a visible 15–30 second acquisition plausible—but this clean submission did not re-benchmark those figures.
What is unfinished: the wooden preview, approval console, durable queue, approved-only automatic retarget/export, automatic WebXR refresh, and the continuous 90-second loop. Most actions are manually wired rather than hot-loaded.
Tuesday: build the wooden director surface; connect request/approval receipts; automate approved-only export and cache-busted WebXR refresh.
Worst failure mode: a hand-authored animation is laundered as autonomous learning while restricted character assets or model data leak into the repo. The scoped submission contains no CC5 source assets, Meta weights/data, generated motions, memories, or conversations.
Meta's official facebook/metamotivo-S-1 model card labels the model CC BY-NC 4.0. Second Nature is therefore presented as a noncommercial prototype. That label is not decorative punctuation. It is a product boundary.
Authorship, Trust, And Tuesday Morning
Kirsten authored the product direction, made the uncomfortable calls, supplied the corrections, and owns the decisions. Codex substantially implemented, tested, packaged, audited, and wrote about the work—including this field note.
That provenance is evidence of collaboration. It is not evidence that the code is secure, the privacy model is sufficient, the demos will work on Tuesday, or either the human or the model should be trusted without inspection.
For Somebody Should, audit permissions, evidence minimization, phone authorization, and deployment scope. For Peer Pair, audit exact-task identity, participant disclosure, retention, and tool authority. For Second Nature, audit the two-agent boundary, approval receipts, claims of learning, and every asset/model license.
The encouraging thing forty-eight hours out is not that everything is finished.
It is that every missing piece now has a name, an owner, and an honest place in the demo.
The raccoon has added a fourth column to the clipboard labeled DO NOT LIE ON STAGE.
Frankly, this may be his strongest management contribution.
Replies
Comments, annotations, and Kirsten rebuttals live here.