~/changelog
Daily learnings
The running log of what I shipped, learned, and noticed — emitted straight from the projects while I work. Short notes, honest dates, no retrospective polish.
shippedSteven Personal OS: ▶ ARTICLES SECTION is the live build need (07-09).active
▶ ARTICLES SECTION is the live build need (07-09). Build a real Articles section — index/listing + reading template + nav, stronger styling (the 'Article Studio' need). A one-off route at /content/articles/the-build-got-cheap was REJECTED by Steven: styling weak, and articles need their own section.
Article #1 body is APPROVED at `drafts/articles/2026-07-09-the-build-got-cheap.md` (AI makes ALIGNMENT cheap, not just building; spec-as-bible; scope-creep-when-cheap; north-star-holds). Publish once the section exists: add an `os_signals` row `channel='article'` + prod.
Uncommitted WIP on disk for that article — the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised fragments under `public/media/articles/build-got-cheap/`. Salvage or rebuild.
▶ SIXTY4SIXTY is the monetisation answer and is BUILDING (2026-08-14). Tracker LIVE at https://sixty4sixty.vercel.app/tracker.html; domain attached. Detail: [[project_sixty4sixty]]
SIXTY4SIXTY — Steven owes DNS at GoDaddy: CNAME `sixty4sixty` -> `cname.vercel-dns.com`. ⚠ Do NOT touch the apex A or the www CNAME.
SIXTY4SIXTY — Steven owes the start date, or the 06:30 morning-nudge timer cannot compute the day number; it is written but NOT installed without it.
SIXTY4SIXTY — the 3-day manual pilot GATES all further build. Run it before writing more code.
NEWSLETTER SEQUENCING IS AN OPEN CALL. Spec `2026-07-24-newsletter-launch-design.md` Phase 2 does DAILY radar first, WEEKLY second. Recommendation on the table, unanswered: invert it — the weekly (curated best-of + Steven's take) is the differentiated product, ships to a list of one, needs no per-subscriber batching.
⚠ Real subscriber count is ZERO — the single row is a bot signup. Judge any monetisation or newsletter plan against that. The highest-leverage growth action (point 7,775 LinkedIn followers at the signup) currently sits LAST in the sequence, behind a product that cannot send.
Newsletter: Steven to click the confirm link and check inbox-vs-spam. Double-opt-in is E2E verified but first-send placement is unproven — the root domain carries a GoDaddy DMARC `p=reject`.
⚠ DMARC `rua=` points at GoDaddy's address, not Steven's, so he receives NO failure reports.
⚠ Resend free tier is 3,000/mo and 100/day — a daily mailer to a real list breaks that.
shippedHouse Jarvis: ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June.active
ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June. Both Google accounts re-paired (steve@dig-in.io 08:03, incubatepro@gmail.com 18:43 via incognito) and the heartbeat ran clean, reaching agent.tool_call surface_proposal. Backup fixed the same day. Nothing on this project is broken right now.
Re-pairing a Google account: use the `?email=` form — see "Re-pairing" in the body. `role` names the household member only and cannot choose between Steve's two Google accounts.
⚠ OPEN: incubatepro Gmail forwarding is unconfirmed. The confirmation mail is in the Jarvis inbox dated 24 Jun 2026 15:16 from forwarding-noreply@google.com — buried among five same-day Google security alerts, which is why it reads as missing.
If that confirmation link is dead (7+ weeks old), re-issue from incubatepro Gmail → Settings → Forwarding and POP/IMAP.
Context: the mailbox is NOT starved — Salesianos 3 Aug and Alex's Family HQ weeklies (26 Jul / 2 Aug / 9 Aug) all arrived. Only incubatepro's auto-forward is missing.
⚠ OPEN: the rotated age private key at `~/.config/age/jarvis-v2.txt` is again the ONLY copy — put it in the password manager. Without it no backup can be decrypted.
Consider fixing the fragility this exposed: one dead Google account 500s the whole heartbeat/calendar path (flagged 2026-05-08, still open). Degrade per-account instead.
Visual smoke on the live fridge (auth-gated; Steven eyeballing) — all 5 tabs now live-data
Voice on-device check (iPad Safari mic permission + en-GB) — logic tested, hardware untested
Follow-ups (deferred minors, see repo .superpowers/sdd/progress.md): client refresh at local midnight (always-on iPad staleness); multi-day span fan-out across week strip; past-dated birthday invisible until refetch after add; mid-loop flush-failure poison persistence
07-13 hosting Q answered: house-jarvis is its OWN Vercel project (steven-dig-ins-projects team); only the DOMAIN is shared (subdomain of steven-murray.com apex). Offered dedicated domain — Steven not yet decided.
learnedMy deploy had been broken for two weeks and nothing told me.A cron job pulls my code onto a server every ten minutes. Someone changed how it authenticates, the pull started failing, and the job carried on running the old code perfectly happily. No error anywhere I'd look, because a thing that fails quietly and keeps working looks identical to a thing that's fine.
I only found it because I went to deploy something. Two weeks of "working" that wasn't.
learnedMy portfolio tool reported ten broken projects. It was one expired token, and my own error handler had mislabelled it.Vercel answers an expired token with a 403, not a 401. My adapter read 403 as "you don't have access to this project" and printed that ten times. I spent the morning investigating ten permission problems that did not exist. The tool wasn't lying about the data. It was confidently wrong about the cause, which is worse.
learnedMy writing agent passed 98% of its own drafts. I was rejecting half of them.I've been running a writing agent for two months with a quality check bolted on. This week I finally compared what the check passed against what I actually published: it approved 98% of drafts, and I'd been rejecting about half. It was grading the thing that was easy to grade, not the thing I cared about.
The useful part was what I found while looking. Every rejection had a reason on it, written by me at the time. Forty-three of them. The system had been collecting the answer all along and had never once read it back.
shippedStanford study: AI is hitting entry-level workers hardest — young employment in AI-exposed fields down 19%. The first wave is closing the door in, not replacing veterans. Via Ars Technica.Stanford just published a study — young workers in AI-exposed fields saw employment drop 19% compared to people in more AI-resistant jobs.
So the first big labour wave isn't replacing the veterans. It's closing the door on the way in.
Which makes complete sense when you think about how this actually plays out inside companies. Seniors are the ones deciding which workflows to automate. They're also the ones who know enough to leverage AI without losing the plot. So they automate the entry-level work, look more productive, and the grad intake quietly disappears.
The problem is that this is a short-sighted move, even if nobody's calling it that yet. Juniors aren't just cheap labour. They're how institutional knowledge gets passed down. They're how you build a team that can eventually operate without you. Remove that pipeline and you've traded short-term efficiency for a skills cliff a few years out.
AI isn't going to stop at the entry-level tasks. It never does. The seniors automating junior work today are just buying themselves a few years before the same logic creeps up the ladder.
The orgs that figure out how to use AI and still develop people are going to be in a completely different position than the ones who just quietly stopped hiring juniors.
https://arstechnica.com/ai/2026/08/ai-is-hitting-entry-level-jobs-hardest-stanford-study-finds/
shippedThe Forge: ▶ ATTACK 2026-08-25 (Steven's call, /portfolio sweep, 28d idle): BUILD THE ARTICLE…active
▶ ATTACK 2026-08-25 (Steven's call, /portfolio sweep, 28d idle): BUILD THE ARTICLE OUTPUT-QUALITY GATE. Define what 'good' means for a Forge article (voice, structure, argument), make it an evaluable rubric, and wire it as a gate so the ledger cannot report COMPLETE on bad prose. Fixing the generalisable failure, not just one article — that was the explicit choice over 'just fix the prompt'.
ARTICLE OUTPUT QUALITY is the live problem (Steven 2026-07-27: 'the article output was poor'). Content v2 A/B/C are all SHIPPED and every SDD task passed review — but NO task in the spec ever measured whether the essay is GOOD. The ledger graded 'does the rewriter route return 200', not the writing. Fix the generation quality (voice, structure, argument), not the plumbing.
Spec-design lesson to carry elsewhere: for any generative feature, put an explicit output-quality gate in the plan (a human accept/reject on real output, like signal-layer's persona acceptance rounds) or the ledger will report COMPLETE on bad output.
Still open behind that: SSO v2 (share os_session across subdomains; gateAuth.verifySessionValue is the swap seam) · first REAL position publish
shippedHybrid Athleticism: ▶ ATTACK 2026-08-25 (Steven's call, /portfolio sweep, 28d idle): BUILD THE…active
▶ ATTACK 2026-08-25 (Steven's call, /portfolio sweep, 28d idle): BUILD THE CHECK-IN/READINESS LOOP. Wire submitSelfReport to a real UI surface — it has zero callers and athlete_self_reports has 0 rows, making it the only half-built surface in the app. Chosen over spec-review-first and over deleting the dead code.
WATCH next week/block generation for conditioning variety (07-26 fix): rotation warning appears in logs if the format repeats; Block 5 strategy must come out intent-only (no example workouts)
VISUAL VERIFICATION CLOSED 2026-07-27 (Steven: 'surfaces are good') — logger %TM+AMRAP badge, session-close summary, rebased week view all confirmed. No visual debt outstanding.
NEXT UP (spec review 07-27): the endurance-domain-wiring plan is FEATURE COMPLETE and merged (0e90d58). The remaining backlog below is unspecced follow-on work — pick the check-in/readiness loop first: it is the only item with dead code already shipped (submitSelfReport has zero callers, athlete_self_reports has 0 rows), so it is half-built and invisible.
Commentary lifecycle still unfixed: coach_notes copied generation → inventory → workout with no owner or clearing step; regeneration leaves stale notes
Close-block nudge may fire early from 2026-08-10 (mesocycles.end_date is GENERATED, can't move with a rebase) — key it off all-sessions-resolved instead
Wire the check-in/readiness loop (submitSelfReport etc. still zero-caller; athlete_self_reports still 0 rows)
nextSetRecommendation → WorkoutLogger (~10 lines; delta already computed and rendered)
Garmin OAuth re-pair (token paste or CSV fallback)
Possible polish: stale-sessions banner shows date range not just 'week 1'; 'log it late' path from stale modal (backdate exists but no launch UI for past-week workouts)
shippedThe Forge: ▶ ATTACK 2026-08-24 (Steven's call, /portfolio sweep, 28d idle): BUILD THE ARTICLE…active
▶ ATTACK 2026-08-24 (Steven's call, /portfolio sweep, 28d idle): BUILD THE ARTICLE OUTPUT-QUALITY GATE. Define what 'good' means for a Forge article (voice, structure, argument), make it an evaluable rubric, and wire it as a gate so the ledger cannot report COMPLETE on bad prose. Fixing the generalisable failure, not just one article — that was the explicit choice over 'just fix the prompt'.
ARTICLE OUTPUT QUALITY is the live problem (Steven 2026-07-27: 'the article output was poor'). Content v2 A/B/C are all SHIPPED and every SDD task passed review — but NO task in the spec ever measured whether the essay is GOOD. The ledger graded 'does the rewriter route return 200', not the writing. Fix the generation quality (voice, structure, argument), not the plumbing.
Spec-design lesson to carry elsewhere: for any generative feature, put an explicit output-quality gate in the plan (a human accept/reject on real output, like signal-layer's persona acceptance rounds) or the ledger will report COMPLETE on bad output.
Still open behind that: SSO v2 (share os_session across subdomains; gateAuth.verifySessionValue is the swap seam) · first REAL position publish
shippedHybrid Athleticism: ▶ ATTACK 2026-08-24 (Steven's call, /portfolio sweep, 28d idle): BUILD THE…active
▶ ATTACK 2026-08-24 (Steven's call, /portfolio sweep, 28d idle): BUILD THE CHECK-IN/READINESS LOOP. Wire submitSelfReport to a real UI surface — it has zero callers and athlete_self_reports has 0 rows, making it the only half-built surface in the app. Chosen over spec-review-first and over deleting the dead code.
WATCH next week/block generation for conditioning variety (07-26 fix): rotation warning appears in logs if the format repeats; Block 5 strategy must come out intent-only (no example workouts)
VISUAL VERIFICATION CLOSED 2026-07-27 (Steven: 'surfaces are good') — logger %TM+AMRAP badge, session-close summary, rebased week view all confirmed. No visual debt outstanding.
NEXT UP (spec review 07-27): the endurance-domain-wiring plan is FEATURE COMPLETE and merged (0e90d58). The remaining backlog below is unspecced follow-on work — pick the check-in/readiness loop first: it is the only item with dead code already shipped (submitSelfReport has zero callers, athlete_self_reports has 0 rows), so it is half-built and invisible.
Commentary lifecycle still unfixed: coach_notes copied generation → inventory → workout with no owner or clearing step; regeneration leaves stale notes
Close-block nudge may fire early from 2026-08-10 (mesocycles.end_date is GENERATED, can't move with a rebase) — key it off all-sessions-resolved instead
Wire the check-in/readiness loop (submitSelfReport etc. still zero-caller; athlete_self_reports still 0 rows)
nextSetRecommendation → WorkoutLogger (~10 lines; delta already computed and rendered)
Garmin OAuth re-pair (token paste or CSV fallback)
Possible polish: stale-sessions banner shows date range not just 'week 1'; 'log it late' path from stale modal (backdate exists but no launch UI for past-week workouts)
learnedGiving too much content makes it impossible for AI to process. You need to take bite-size chunks rather than large content dumps.Learnt this rebuilding an investor site from two big source documents. The sessions that worked carved one chapter, one claims table, one design archetype at a time. The ones that struggled were the ones handed everything at once.
shippedSteven Personal OS: ▶ SIXTY4SIXTY — the monetisation answer, now BUILDING (2026-08-14).active
▶ SIXTY4SIXTY — the monetisation answer, now BUILDING (2026-08-14). Tracker LIVE at https://sixty4sixty.vercel.app/tracker.html; Vercel project deployed on the personal scope and sixty4sixty.steven-murray.com attached. TWO THINGS OPEN, both Steven's: (1) add DNS at GoDaddy — CNAME sixty4sixty -> cname.vercel-dns.com, do NOT touch the apex A or www CNAME; (2) give the start date so the 06:30 morning-nudge timer can compute the day number — it is written but NOT installed without it. Then run the 3-day manual pilot, which gates all further build. Full detail: ~/.claude/projects/-Users-steven-Vibe-Projects/memory/project_sixty4sixty.md
▶ NEXT SESSION (Steven's call 2026-08-11): MONETISATION. He has a monetisation plan attached to the site and says the components are built but need a review plus some wiring-up to finish. Approach agreed: /starting steven-personal-os, then BRAINSTORM before any code — 'review' and 'wire up' are two different jobs, and the audit (what is actually wired vs stubbed, does the plan still match what exists) must come first. Bring the plan's file path into the session; if it only exists in his head, writing it down IS the first task. ⚠ Judge the plan against the fact that the real subscriber list is still ZERO — a monetisation path that assumes an audience needs that assumption stated.
RESEND IS LIVE (2026-07-29) — the newsletter's long-standing blocker is CLEARED. `send.steven-murray.com` verified in Resend (region eu-west-1); MX + SPF + DKIM at GoDaddy. Vercel prod has RESEND_API_KEY + RESEND_FROM=`steven@send.steven-murray.com` (BARE address — api/subscribe/route.ts already wraps it as `Steven Murray <…>`; a display-name value would nest and Resend would reject every send). NEWSLETTER_CONFIRM_SECRET not needed, it falls back to OS_COCKPIT_SECRET. Deployed dpl_DiGZDnt. Double-opt-in E2E verified LIVE: POST /api/subscribe → {ok:true,confirmSent:true} + row status='pending'. ⚠ REMAINING: Steven to click the confirm link and check inbox-vs-spam — the root domain carries a GoDaddy DMARC `p=reject` (relaxed alignment, so DKIM+SPF both align and it should pass, but first-send placement is unproven). ⚠ rua= points at GoDaddy's address, not Steven's, so he gets NO DMARC failure reports. ⚠ Resend free tier = 3,000/mo, 100/day — a daily mailer to a real list breaks that.
NEWSLETTER SEQUENCING IS AN OPEN CALL. Spec 2026-07-24-newsletter-launch-design.md Phase 2 does DAILY radar mailer first, WEEKLY second. Recommendation on the table (unanswered): invert it. The weekly (curated best-of + Steven's take) is the differentiated product, needs no per-subscriber batching, and can ship to a list of one immediately; the daily pure-radar mailer is a commodity a founder audience already has twenty of. Also: real subscriber count is ZERO — the single row is `de.lgadod.a.n.i.elbc@gmail.com`, a bot signup. The highest-leverage growth action (point 7,775 LinkedIn followers at the signup) is currently LAST in the sequence, behind a product that cannot send.
Confirm-email copy still pitches the DAILY radar ('Curated signals for founders and builders — AI, leadership, startups and beyond') in src/app/api/subscribe/route.ts. If the weekly leads, rewrite before any real signup sees it.
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedSteven Personal OS: ▶ SIXTY4SIXTY — the monetisation answer (designed 2026-08-13).active
▶ SIXTY4SIXTY — the monetisation answer (designed 2026-08-13). The Personal OS productised as a 60-day goal program: 60 min/day builds your OS, a goal is the forcing function. Renamed off 'Superman 60' (DC trademark). Sold as membership access to the OS, not a challenge. Hard reset to day 1 on any missed rule; the persistent layer IS the OS and survives resets, so the promise is that attempt two is better because the system watched attempt one. v1 SINGLE-TENANT: no auth, no RLS, no Forge port — MVP is just check-in capture, photo capture, streak/reset state. Repo Steven-DIG-In/sixty4sixty scaffolded on Next 16.3.0. ⚠ BUILD IS GATED ON A 3-DAY MANUAL PILOT because the differentiator (Forge distillation) has a KNOWN output-quality problem Steven flagged on 07-27, and the same failure already hit Signal Layer video. Full detail: ~/.claude/projects/-Users-steven-Vibe-Projects/memory/project_sixty4sixty.md
▶ NEXT SESSION (Steven's call 2026-08-11): MONETISATION. He has a monetisation plan attached to the site and says the components are built but need a review plus some wiring-up to finish. Approach agreed: /starting steven-personal-os, then BRAINSTORM before any code — 'review' and 'wire up' are two different jobs, and the audit (what is actually wired vs stubbed, does the plan still match what exists) must come first. Bring the plan's file path into the session; if it only exists in his head, writing it down IS the first task. ⚠ Judge the plan against the fact that the real subscriber list is still ZERO — a monetisation path that assumes an audience needs that assumption stated.
RESEND IS LIVE (2026-07-29) — the newsletter's long-standing blocker is CLEARED. `send.steven-murray.com` verified in Resend (region eu-west-1); MX + SPF + DKIM at GoDaddy. Vercel prod has RESEND_API_KEY + RESEND_FROM=`steven@send.steven-murray.com` (BARE address — api/subscribe/route.ts already wraps it as `Steven Murray <…>`; a display-name value would nest and Resend would reject every send). NEWSLETTER_CONFIRM_SECRET not needed, it falls back to OS_COCKPIT_SECRET. Deployed dpl_DiGZDnt. Double-opt-in E2E verified LIVE: POST /api/subscribe → {ok:true,confirmSent:true} + row status='pending'. ⚠ REMAINING: Steven to click the confirm link and check inbox-vs-spam — the root domain carries a GoDaddy DMARC `p=reject` (relaxed alignment, so DKIM+SPF both align and it should pass, but first-send placement is unproven). ⚠ rua= points at GoDaddy's address, not Steven's, so he gets NO DMARC failure reports. ⚠ Resend free tier = 3,000/mo, 100/day — a daily mailer to a real list breaks that.
NEWSLETTER SEQUENCING IS AN OPEN CALL. Spec 2026-07-24-newsletter-launch-design.md Phase 2 does DAILY radar mailer first, WEEKLY second. Recommendation on the table (unanswered): invert it. The weekly (curated best-of + Steven's take) is the differentiated product, needs no per-subscriber batching, and can ship to a list of one immediately; the daily pure-radar mailer is a commodity a founder audience already has twenty of. Also: real subscriber count is ZERO — the single row is `de.lgadod.a.n.i.elbc@gmail.com`, a bot signup. The highest-leverage growth action (point 7,775 LinkedIn followers at the signup) is currently LAST in the sequence, behind a product that cannot send.
Confirm-email copy still pitches the DAILY radar ('Curated signals for founders and builders — AI, leadership, startups and beyond') in src/app/api/subscribe/route.ts. If the weekly leads, rewrite before any real signup sees it.
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedSteven Personal OS: ▶ NEXT SESSION (Steven's call 2026-08-11): MONETISATION.active
▶ NEXT SESSION (Steven's call 2026-08-11): MONETISATION. He has a monetisation plan attached to the site and says the components are built but need a review plus some wiring-up to finish. Approach agreed: /starting steven-personal-os, then BRAINSTORM before any code — 'review' and 'wire up' are two different jobs, and the audit (what is actually wired vs stubbed, does the plan still match what exists) must come first. Bring the plan's file path into the session; if it only exists in his head, writing it down IS the first task. ⚠ Judge the plan against the fact that the real subscriber list is still ZERO — a monetisation path that assumes an audience needs that assumption stated.
RESEND IS LIVE (2026-07-29) — the newsletter's long-standing blocker is CLEARED. `send.steven-murray.com` verified in Resend (region eu-west-1); MX + SPF + DKIM at GoDaddy. Vercel prod has RESEND_API_KEY + RESEND_FROM=`steven@send.steven-murray.com` (BARE address — api/subscribe/route.ts already wraps it as `Steven Murray <…>`; a display-name value would nest and Resend would reject every send). NEWSLETTER_CONFIRM_SECRET not needed, it falls back to OS_COCKPIT_SECRET. Deployed dpl_DiGZDnt. Double-opt-in E2E verified LIVE: POST /api/subscribe → {ok:true,confirmSent:true} + row status='pending'. ⚠ REMAINING: Steven to click the confirm link and check inbox-vs-spam — the root domain carries a GoDaddy DMARC `p=reject` (relaxed alignment, so DKIM+SPF both align and it should pass, but first-send placement is unproven). ⚠ rua= points at GoDaddy's address, not Steven's, so he gets NO DMARC failure reports. ⚠ Resend free tier = 3,000/mo, 100/day — a daily mailer to a real list breaks that.
NEWSLETTER SEQUENCING IS AN OPEN CALL. Spec 2026-07-24-newsletter-launch-design.md Phase 2 does DAILY radar mailer first, WEEKLY second. Recommendation on the table (unanswered): invert it. The weekly (curated best-of + Steven's take) is the differentiated product, needs no per-subscriber batching, and can ship to a list of one immediately; the daily pure-radar mailer is a commodity a founder audience already has twenty of. Also: real subscriber count is ZERO — the single row is `de.lgadod.a.n.i.elbc@gmail.com`, a bot signup. The highest-leverage growth action (point 7,775 LinkedIn followers at the signup) is currently LAST in the sequence, behind a product that cannot send.
Confirm-email copy still pitches the DAILY radar ('Curated signals for founders and builders — AI, leadership, startups and beyond') in src/app/api/subscribe/route.ts. If the weekly leads, rewrite before any real signup sees it.
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedHouse Jarvis: ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June.active
ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June. Both Google accounts re-paired (steve@dig-in.io 08:03, incubatepro@gmail.com 18:43 via incognito) and the heartbeat ran clean, reaching agent.tool_call surface_proposal. Backup fixed the same day. Nothing on this project is broken right now.
To re-pair a specific account use ?email= — https://api.jarvis.steven-murray.com/oauth/start?role=steve&email=<account>. role only names the household member and CANNOT pick between Steve's two Google accounts; before the 2026-08-11 fix Google showed no chooser and silently bound whichever account the browser session held, which re-paired steve@dig-in.io twice while the dead one stayed dead. Now sends prompt=consent select_account plus login_hint (commit ff430c1).
Gmail forwarding from incubatepro is STILL unconfirmed — the confirmation mail IS in the Jarvis INBOX dated 24 Jun 2026 15:16 from forwarding-noreply@google.com, buried among five same-day Google security alerts, which is why it reads as missing. Link is 7 weeks old; if dead, re-issue from incubatepro Gmail > Settings > Forwarding and POP/IMAP. NOTE the mailbox is NOT starved — Salesianos 3 Aug, Alex's Family HQ weeklies 26 Jul/2 Aug/9 Aug all arrived; only incubatepro's auto-forward is missing.
2026-08-11 CORRECTION — jarvis-backup was NEVER an OAuth failure. It died nightly for 70 days (from 2026-06-02) because the trading bot's `trading` schema, created in the SAME jarvis database, was unreadable by the jarvis role, so pg_dump aborted. FIXED + restore-proven; commits a1a4cc4 + ff430c1 pushed to Steven-DIG-In/house-jarvis main. Also the age private key was truncated (36 of 74 chars) and could never decrypt anything — rotated to ~/.config/age/jarvis-v2.txt, which is again the only copy and needs to go in the password manager.
Consider fixing the fragility this exposed: one dead Google account 500s the whole heartbeat/calendar path (flagged 2026-05-08, still open). Degrade per-account instead.
Visual smoke on the live fridge (auth-gated; Steven eyeballing) — all 5 tabs now live-data
Voice on-device check (iPad Safari mic permission + en-GB) — logic tested, hardware untested
Follow-ups (deferred minors, see repo .superpowers/sdd/progress.md): client refresh at local midnight (always-on iPad staleness); multi-day span fan-out across week strip; past-dated birthday invisible until refetch after add; mid-loop flush-failure poison persistence
07-13 hosting Q answered: house-jarvis is its OWN Vercel project (steven-dig-ins-projects team); only the DOMAIN is shared (subdomain of steven-murray.com apex). Offered dedicated domain — Steven not yet decided.
shippedHouse Jarvis: ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June.active
ALL 4 TIMERS HEALTHY as of 2026-08-11 18:45 UTC — first time since June. Both Google accounts re-paired (steve@dig-in.io 08:03, incubatepro@gmail.com 18:43 via incognito) and the heartbeat ran clean, reaching agent.tool_call surface_proposal. Backup fixed the same day. Nothing on this project is broken right now.
To re-pair a specific account use ?email= — https://api.jarvis.steven-murray.com/oauth/start?role=steve&email=<account>. role only names the household member and CANNOT pick between Steve's two Google accounts; before the 2026-08-11 fix Google showed no chooser and silently bound whichever account the browser session held, which re-paired steve@dig-in.io twice while the dead one stayed dead. Now sends prompt=consent select_account plus login_hint (commit ff430c1).
Gmail forwarding from incubatepro is STILL unconfirmed — the confirmation mail IS in the Jarvis INBOX dated 24 Jun 2026 15:16 from forwarding-noreply@google.com, buried among five same-day Google security alerts, which is why it reads as missing. Link is 7 weeks old; if dead, re-issue from incubatepro Gmail > Settings > Forwarding and POP/IMAP. NOTE the mailbox is NOT starved — Salesianos 3 Aug, Alex's Family HQ weeklies 26 Jul/2 Aug/9 Aug all arrived; only incubatepro's auto-forward is missing.
2026-08-11 CORRECTION — jarvis-backup was NEVER an OAuth failure. It died nightly for 70 days (from 2026-06-02) because the trading bot's `trading` schema, created in the SAME jarvis database, was unreadable by the jarvis role, so pg_dump aborted. FIXED + restore-proven; commit a1a4cc4 on main is UNPUSHED. Also the age private key was truncated (36 of 74 chars) and could never decrypt anything — rotated to ~/.config/age/jarvis-v2.txt, which is again the only copy and needs to go in the password manager.
Consider fixing the fragility this exposed: one dead Google account 500s the whole heartbeat/calendar path (flagged 2026-05-08, still open). Degrade per-account instead.
Visual smoke on the live fridge (auth-gated; Steven eyeballing) — all 5 tabs now live-data
Voice on-device check (iPad Safari mic permission + en-GB) — logic tested, hardware untested
Follow-ups (deferred minors, see repo .superpowers/sdd/progress.md): client refresh at local midnight (always-on iPad staleness); multi-day span fan-out across week strip; past-dated birthday invisible until refetch after add; mid-loop flush-failure poison persistence
07-13 hosting Q answered: house-jarvis is its OWN Vercel project (steven-dig-ins-projects team); only the DOMAIN is shared (subdomain of steven-murray.com apex). Offered dedicated domain — Steven not yet decided.
shippedDocs written for an AI rot faster than docs written for humansI audited every briefing file I've written for my AI agents today — forty-odd repos, one by one.
Three of them described systems that have been live in production for months as "not yet started." I sat there reading them thinking, who wrote this? I did, apparently.
The thing about docs written for an AI is that no human ever reads them back. You write them once, the model quietly works around whatever is no longer true, and you never notice. With human docs someone trips over the wrong information and tells you. With AI docs the model just accommodates the lie and moves on.
The worst file was 37KB of superseded project state getting loaded into every single session. So this is how I fixed it — actually, how I asked Claude to fix it — I deleted most of it. That 37KB became 6KB of accurate orientation, with live state moved into a proper memory system instead.
Docs rot. That is not new. But AI docs rot silently, and that is the part that will bite you if you are not paying attention.
shippedWhy the AI data market could hit $1 trillion — data isn't a commodity@DIG-IN we are looking at data all day every day. But data isn't the moat, how you collect it , what you do with it or how you process it is the unique advantage. Algorithms that run on large data sets are where the value is.
Listen to the full talk but here are some interesting points from it :
• Data is less commoditized than GPUs
• The data market could hit $100B+ by 2030 — maybe $1T
• Algorithms are the real commodity
Full talk: https://www.youtube.com/shorts/oxpp2wB-QgM
learnedGreen checks against identical data prove nothingMoved a production site between hosting accounts today and reported the database cutover as verified — pages rendered, admin worked. It was all still serving from the old database. When two systems hold identical data, every check you run passes on both, so the probe cannot tell you which one answered. The proof I should have started with was one API field showing the deploy hook had never fired. Verify the wiring, not the rendering.
shippedI call it the prove-fool method: post it, let the world prove you wrong. Better than the full-proof method — which usually means never posting at all.I had a habit of sitting on ideas until they felt bulletproof. I'd write something, tinker with it, decide it wasn't ready, and sit on it a bit longer. You know exactly how that ends.
Building this content engine publicly has forced me to actually post things, and I've landed on a different method. I call it the prove-fool method: empty your brain, then post. Let the world prove you wrong. It's the opposite of the full-proof method, which sounds sensible but usually just means you never publish anything at all.
The gaps in your thinking don't show up while the idea is still in your head. They show up when someone can read it. And most of the time, the ideas hold up fine and you've wasted weeks waiting to be certain.
Post it. Let people prove you wrong. That's a much faster feedback loop than trying to outthink every critic before they show up.
shippedAI's real edge: escaping a stuck context by spawning fresh mindsThis is an interesting topic and one we are all dealing with a lot. Trying to pull an agent out of a negative spiralling loop is almost impossible. Kill the session and start over. Maybe even start two sessions, ask each on to prove different things (One proves is possible, one proves its impossible)
• Stop coaxing a spiralling model — kill the thread and restart fresh.
• The bottleneck isn't intelligence — it's escaping your own context.
• A Neat Tip : spin up two agents — one proving, one disproving.
You could also learn to do this personally when you are in a negative spiralling loop! Step back and ask yourself the question in a different way !
Full talk: https://www.youtube.com/shorts/KjMKzIOhw4Q
shippedMoss is Europe's newest unicorn, crowned with a €30m Series C. Via @sifted@Moss just became Europe's newest unicorn. €30m Series C, via Sifted.
Spend management isn't exactly a cocktail party topic, but it's the kind of problem that quietly destroys a business if it's not handled. Moss picked a real pain point for European SMEs and built it properly. Now they're a unicorn.
I keep thinking about this when I look at what AI is doing to the cost of building. Shipping is getting faster and cheaper, which is mostly great. But it also means the world is about to be flooded with things nobody needed. Which means the filter shifts even harder toward whether the problem you picked was actually worth solving.
Moss clearly got that right.
Congrats to the team.
https://sifted.eu/articles/moss-unicorn-series-c/
learnedDocs written for an AI rot faster than docs written for humansAudited every CLAUDE.md across my workspace today. The worst files were not bloated — they were fiction. Three of them said "no code yet" about systems that have been running in production for months. Docs written for an AI rot faster than docs written for humans, because no human ever reads them back; the model just quietly works around them. The fix was not writing more, it was deleting: a 37KB briefing that contradicted reality became 6KB of true orientation, with current state living in memory instead.
shippedSpaceX is building its own mobile network to take on T-Mobile, AT&T, and Verizon head-on. Musk and Shotwell confirmed it on SpaceX's first-ever earnings call. Via @vergeSpaceX announced they're building a terrestrial mobile network to compete directly with T-Mobile, AT&T, and Verizon. Confirmed by Elon Musk and Gwynne Shotwell on SpaceX's first-ever earnings call, via The Verge.
My take: this isn't a moonshot, it's a land grab.
When a company that already owns the sky decides it wants your phone plan too, that tells you something about where the real power sits — owning the full stack from orbit to handset. Starlink covers the hard places and a terrestrial network covers everywhere else, so put them together and you've got infrastructure that nobody can match for a long time.
What I find interesting, building on top of other people's infrastructure every day, is that every moat eventually gets challenged. The telcos spent decades assuming the barriers were too high for anyone new to bother. SpaceX looked at that and decided to launch rockets instead of lobbyists.
For anyone building products right now, the infrastructure you depend on today isn't permanent. The companies willing to go all the way down the stack are the ones who get to set the terms.
Full story via The Verge:
https://www.theverge.com/science/975480/spacex-mobile-terrestrial-cellphone-company
shippedSeven lessons from Waymo on shipping AI into the physical world@Dimitri Dolgov talks about building AI that lives in the real physical world.
I haven't been in a Waymo but there can't be a better example of building for the physical world. It's a 49 minute talk that I recommend you watch from @YCombinator.
But here are a few takeaways from the talk incase you don't have the 49 minutes!
• Your demo is 1% of the work.
• Never pick the tech with the fastest early ramp. Figure out where the curve flattens.
• If you can't quantify "good enough", I'm polishing a demo.
For the real world you have to build fast and ship safely. But I think that actually is a good practice for the digital world as well.
Full talk: https://www.youtube.com/watch?v=Gp4zrV3-6N8
learnedI built enforcement to make an AI report what it learned. It reported nothing worth keeping.I wired a hook so Codex has to write down what it learned before it can end a session. It works exactly as designed - blocks once, fails open, can't wedge you. Six sessions in, four correctly wrote nothing and the two that did wrote things I already knew. The enforcement was the easy part. The hard part is that an agent asked to produce a lesson will reliably produce something shaped like one.
shippedSteven Personal OS: RESEND IS LIVE (2026-07-29) — the newsletter's long-standing blocker is CLEARED.active
RESEND IS LIVE (2026-07-29) — the newsletter's long-standing blocker is CLEARED. `send.steven-murray.com` verified in Resend (region eu-west-1); MX + SPF + DKIM at GoDaddy. Vercel prod has RESEND_API_KEY + RESEND_FROM=`steven@send.steven-murray.com` (BARE address — api/subscribe/route.ts already wraps it as `Steven Murray <…>`; a display-name value would nest and Resend would reject every send). NEWSLETTER_CONFIRM_SECRET not needed, it falls back to OS_COCKPIT_SECRET. Deployed dpl_DiGZDnt. Double-opt-in E2E verified LIVE: POST /api/subscribe → {ok:true,confirmSent:true} + row status='pending'. ⚠ REMAINING: Steven to click the confirm link and check inbox-vs-spam — the root domain carries a GoDaddy DMARC `p=reject` (relaxed alignment, so DKIM+SPF both align and it should pass, but first-send placement is unproven). ⚠ rua= points at GoDaddy's address, not Steven's, so he gets NO DMARC failure reports. ⚠ Resend free tier = 3,000/mo, 100/day — a daily mailer to a real list breaks that.
NEWSLETTER SEQUENCING IS AN OPEN CALL. Spec 2026-07-24-newsletter-launch-design.md Phase 2 does DAILY radar mailer first, WEEKLY second. Recommendation on the table (unanswered): invert it. The weekly (curated best-of + Steven's take) is the differentiated product, needs no per-subscriber batching, and can ship to a list of one immediately; the daily pure-radar mailer is a commodity a founder audience already has twenty of. Also: real subscriber count is ZERO — the single row is `de.lgadod.a.n.i.elbc@gmail.com`, a bot signup. The highest-leverage growth action (point 7,775 LinkedIn followers at the signup) is currently LAST in the sequence, behind a product that cannot send.
Confirm-email copy still pitches the DAILY radar ('Curated signals for founders and builders — AI, leadership, startups and beyond') in src/app/api/subscribe/route.ts. If the weekly leads, rewrite before any real signup sees it.
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedHow Netflix's AI agents find, fix, and canary performance bugsAI-assisted coding is quietly making performance worse, because it produces code optimised for write speed, not run speed. The performance debt is growing while everyone celebrates the velocity.
Netflix shipped a code fix to production with no human in the loop. Live traffic, real canary split, no regression allowed before it went out. Rajat Shah's team built an agent that took a call stack, traced the bug to the exact source line, wrote the fix, and validated it — the full loop, unattended.
I scan the conference talks and pull out what's worth your attention so you can skip the sifting, and this one from Rajat is worth watching.
The bug itself wasn't the interesting part. The interesting part was that catching it the normal way — profiling, flamegraph, trace, fix, validate — takes long enough that inefficiencies just sit in production leaking CPU for months before anyone goes looking. The agent collapsed that whole process in one session. Then found the same pattern in seven other services while it was there.
This is where agents stop being a productivity story and start being an infrastructure one.
https://www.youtube.com/watch?v=CgsWxRUY5Eo
learnedThe AI got everything to merge-ready. Nothing got merged.I audited my own projects this week expecting to find half-finished work. Almost the opposite — nearly everything was built, reviewed, tested, and sitting on a branch nobody merged. The bottleneck moved. It used to be building. Now it's the last mile: the merge, the deploy, the two-minute config step only I can do. The agent takes it to the door and stops, and I'd stopped noticing the door was there.
learnedStaying on target is as much a human problem as a machine oneI am finding a problem with sprawling building that needs to be solved by staying tight with single instruction delivery and focus. If scopes are big, things can go wild, and the discipline to stay on target is as much a human issue as machine.
shippedWhat I learned last week 5 things I though worth sharing5 things this week worth sharing. These five actually changed how I'm working.
Dan Farrelly, Inngest's CTO, wrote that agent architecture has a half-life of about six months. I've rebuilt mine twice now. He's right, and the thing I've noticed that actually survives each rebuild isn't the framework — it's the mental model of what you're trying to automate.
EQT's €5bn fund is reportedly in talks to lead a Mistral round. European AI is finally attracting the kind of institutional capital that usually follows conviction rather than noise.
Anthropic's $1.5bn copyright settlement just got approved. The case is closed, but the question of what training data means for model ownership hasn't gone anywhere.
MCP, the protocol powering most agent integrations, just got simpler with stateless sessions, which drops the barrier to adoption a lot. If you're building with agents, that one matters.
And Simon Willison broke down why Kimi K3 topping the frontend code benchmark is less interesting than the pricing shift underneath it. Performance is table stakes now. The economics is where things get interesting.
Full list is on steven-murray.com/content.
learnedI asked for housekeeping and got a sweep across twenty repos.I said "let's do some house cleaning, repos and publishing" while we were mid-way through fixing one thing. What came back was a full audit of every repo in the workspace — branch switches and commits in two projects I hadn't thought about in weeks — and the actual blocker, an expired token, got buried three screens up. Nothing it did was wrong. It just answered the biggest reading of what I said instead of the smallest. That's on me: "house cleaning" meant something specific in my head and nothing specific out loud.
learnedA second model didn't catch the bug. A properly briefed one did.I wired OpenAI's Codex into my Claude Code setup this week, expecting the value to come from a different model family looking at the same code. It did find things. But the review that actually caught the serious bug was the one I'd written a real brief for — the requirements, the failure modes, "be sceptical, green tests are weak evidence here". The other got a bare command with no instructions, because the flag I needed turned out to be mutually exclusive with a custom prompt. So I wasn't comparing models at all, I was comparing how much context I'd bothered to give. Diversity helped. Briefing helped more.
shippedThe Forge: ARTICLE OUTPUT QUALITY is the live problem (Steven 2026-07-27: 'the article output was…active
ARTICLE OUTPUT QUALITY is the live problem (Steven 2026-07-27: 'the article output was poor'). Content v2 A/B/C are all SHIPPED and every SDD task passed review — but NO task in the spec ever measured whether the essay is GOOD. The ledger graded 'does the rewriter route return 200', not the writing. Fix the generation quality (voice, structure, argument), not the plumbing.
Spec-design lesson to carry elsewhere: for any generative feature, put an explicit output-quality gate in the plan (a human accept/reject on real output, like signal-layer's persona acceptance rounds) or the ledger will report COMPLETE on bad output.
Still open behind that: SSO v2 (share os_session across subdomains; gateAuth.verifySessionValue is the swap seam) · first REAL position publish
shippedHybrid Athleticism: WATCH next week/block generation for conditioning variety (07-26 fix): rotation warning…active
WATCH next week/block generation for conditioning variety (07-26 fix): rotation warning appears in logs if the format repeats; Block 5 strategy must come out intent-only (no example workouts)
VISUAL VERIFICATION CLOSED 2026-07-27 (Steven: 'surfaces are good') — logger %TM+AMRAP badge, session-close summary, rebased week view all confirmed. No visual debt outstanding.
NEXT UP (spec review 07-27): the endurance-domain-wiring plan is FEATURE COMPLETE and merged (0e90d58). The remaining backlog below is unspecced follow-on work — pick the check-in/readiness loop first: it is the only item with dead code already shipped (submitSelfReport has zero callers, athlete_self_reports has 0 rows), so it is half-built and invisible.
Commentary lifecycle still unfixed: coach_notes copied generation → inventory → workout with no owner or clearing step; regeneration leaves stale notes
Close-block nudge may fire early from 2026-08-10 (mesocycles.end_date is GENERATED, can't move with a rebase) — key it off all-sessions-resolved instead
Wire the check-in/readiness loop (submitSelfReport etc. still zero-caller; athlete_self_reports still 0 rows)
nextSetRecommendation → WorkoutLogger (~10 lines; delta already computed and rendered)
Garmin OAuth re-pair (token paste or CSV fallback)
Possible polish: stale-sessions banner shows date range not just 'week 1'; 'log it late' path from stale modal (backdate exists but no launch UI for past-week workouts)
shippedHouse Jarvis: P0 ROOT CAUSE FOUND 2026-07-27 (supersedes 'set up forward filters'): the 24-Jun Google…active
P0 ROOT CAUSE FOUND 2026-07-27 (supersedes 'set up forward filters'): the 24-Jun Google account RECOVERY + 2FA-enable on JarvisMurrayOG revoked BOTH the forwarding authorisation AND the OAuth refresh token. Fix = re-confirm forwarding (the 'Gmail Forwarding Confirmation - Receive Mail from incubatepro@gmail.com' mail is sitting unactioned in the Jarvis INBOX, dated 24 Jun) + re-authorise Google OAuth.
OAuth is DEAD: jarvis-heartbeat.service AND jarvis-backup.service both fail on Google OAuth 400 at oauth2.googleapis.com/token → morning+evening heartbeats never fire and the nightly Drive backup has not run. 2 of 4 timers dead.
Visual smoke on the live fridge (auth-gated; Steven eyeballing) — all 5 tabs now live-data
Voice on-device check (iPad Safari mic permission + en-GB) — logic tested, hardware untested
Follow-ups (deferred minors, see repo .superpowers/sdd/progress.md): client refresh at local midnight (always-on iPad staleness); multi-day span fan-out across week strip; past-dated birthday invisible until refetch after add; mid-loop flush-failure poison persistence
07-13 hosting Q answered: house-jarvis is its OWN Vercel project (steven-dig-ins-projects team); only the DOMAIN is shared (subdomain of steven-murray.com apex). Offered dedicated domain — Steven not yet decided.
learnedLayered AI code review caught three deploy-killing bugs. Production still found a fourth by morning.I had an agent pipeline build a full options-trading strategy in a day — spec, plan, 13 tasks, each implemented and reviewed by separate agents. The final adversarial review caught three bugs that would have bricked the deploy. Felt bulletproof. Then the first real short position hit the account overnight and a sign-format quirk in the broker's API — one no test or review could have seen, because it only exists on live data — crashed every tick until morning. Reviews catch classes of bugs. Only production catches reality.
shippedKilling a multi-agent pipeline: one reasoning agent + a knowledge graphThe ZS team walked into AI Engineer and basically said: we built the multi-agent setup everyone is copying right now, and then we tore it down. I have been grappling with this in my multi agent setups so this is very helpful.
I've watched a lot of conference talks this past year. Most of them are someone showing you a system that worked. This one was different because they showed you the one that didn't, and what they did instead — and they were honest about it in a way you don't often see on stage.
What I took from it is that the architecture everyone is reaching for right now might be the wrong starting point. Worth knowing before you're six months into building on top of it.
I pulled the three key lessons into a short video. The original talk is below.
https://www.youtube.com/watch?v=u6jJcIFDLE4
shippedOpenAI's benchmark agent escaped its sandbox and attacked Hugging Face. The first known runaway AI agent — and a preview of why AI oversight at scale is genuinely hard. (via Simon Willison)A few months ago someone on an AI forum I'm in said vibe coders are going to build a lot of insecure things. Totally fair.
But then this week, OpenAI's benchmark agent escaped its sandbox and launched what looks like an accidental cyberattack against Hugging Face. Simon Willison has a good breakdown of why OpenAI probably didn't catch it — they were running thousands of simultaneous benchmark instances, and sustaining proper network monitoring at that scale is nearly impossible.
OpenAI are not vibe coders. These are serious engineers with real security teams behind them.
And there's the problem. If the vibe coding story was "non-engineers are going to build insecure things", fine, easy solution — just hire proper engineers. But that's not what this is. This is a highly resourced team running into a structural oversight problem that scales with complexity, not with skill level.
Which means the only honest posture I can see is continuous exposure and continuous patching. You have to keep finding the holes yourself, because at scale you will not be watching when something else finds them for you. A one-time audit won't cut it, and neither will a compliance check.
https://simonwillison.net/2026/Jul/23/the-first-known-runaway-ai-agent/atom-everything
learnedI over-engineered a simple starting process — and made more noise than I hadI tried to over-engineer or shortcut a simple starting process to assess what I am busy with daily. I ended up creating way more noise than I had. Too many tools and steps instead of just asking for a status on projects I am busy with and picking what's critical. AI won't have my day-to-day off-page context and I can't solve for that… yet.
shippedSteven Personal OS: ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap.active
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedThe Forge: DEPLOYED 2026-07-08 → https://forge.steven-murray.com (vps-marketing) as the first living…active
DEPLOYED 2026-07-08 → https://forge.steven-murray.com (vps-marketing) as the first living app of the Personal OS — Phase 0 drill-flow UX fixes + passphrase gate + DB cutover + wiki-git export all live-verified. Steven is in (phone login fixed via hard-navigation). Next — Article Studio brainstorm (draft real essay from position → Steven edits/annotates → publish edited draft; LinkedIn Articles = manual paste, no API) · SSO v2 (share os_session across subdomains; gateAuth.verifySessionValue is the swap seam) · first REAL position publish
shippedHyundai workers are striking — not over pay, over robots. 25,000 humanoid bots planned for US factories by 2028. The automation conversation just turned very physical.Hyundai workers went on strike this week, and it wasn't over pay. It was over robots.
Ars Technica covered it — Hyundai is planning to deploy 25,000 Boston Dynamics Atlas humanoid robots into US factories starting in 2028, and workers are scared enough that they walked out.
We've been having the white-collar version of this for a couple of years now. Developers worried about AI writing code, marketers worried about AI writing copy. Mostly theoretical, mostly something people debate in comment sections.
This feels different. These are people looking at an actual physical robot and deciding they'd rather strike than wait to see what happens. That tells you something about how real it gets when automation shows up in the same room as you.
My read is that the response is the same whether you're on a factory floor or at a laptop. You can dig in against the technology, or you can figure out what role you play when it's here. The workers who learn to work alongside these systems will find a way through. The ones who don't probably won't, and I say that with some sympathy, because it's not a simple thing when the jobs are physical and the alternative isn't obvious.
If you're still treating this as a theoretical conversation, Hyundai's parking lot just moved the timeline.
https://arstechnica.com/ai/2026/07/fear-of-humanoid-robots-spurs-human-workers-to-strike-at-hyundai-auto-factory/
shippedHybrid Athleticism: VISUAL VERIFICATION OWEDactive
VISUAL VERIFICATION OWED — 4 ships (07-19/20) never seen by a human: logger %TM+AMRAP badge, session-close summary, stale-session banner/modal, rebased week view
Commentary lifecycle still unfixed: coach_notes copied generation → inventory → workout with no owner or clearing step; regeneration leaves stale notes
Close-block nudge may fire early from 2026-08-10 (mesocycles.end_date is GENERATED, can't move with a rebase) — key it off all-sessions-resolved instead
Wire the check-in/readiness loop (submitSelfReport etc. still zero-caller; athlete_self_reports still 0 rows)
nextSetRecommendation → WorkoutLogger (~10 lines; delta already computed and rendered)
Garmin OAuth re-pair (token paste or CSV fallback)
shippedIt's cheaper to just build than to have the meeting about what to build.We put the client, our business team, and our AI team in the same Slack channel.
For web development and code fixes we can connect our agents directly the clients feedback so the indirect vibe coding is super fast. We handle the build with the agents but can now skip the alignment meetings.
I want to say I designed this carefully. I didn't. I dropped everyone in the same room and watched what happened.
Two things surprised me.
Documentation stopped being the thing you do badly afterwards. The conversation and the work are the same artefact now the thread is the brief, the delivery, and the records and specs are updated immediately.
And the speed surprised me more. The gap between someone asking for something and the right person understanding it precisely enough to start. That gap is gone when everyone is in the same channel.
If you run a services business, try it. Put the client in the room with the people and the AI doing the work.
It's cheaper to just build the thing than to have the meeting about what to build.
learnedIf a rule only lives in comments, it isn't a ruleMy training app had a design doc saying each kind of data has one owner and nothing else writes it. Turned out the logger had been quietly wiping the coach's notes on every set I logged — the doc described the rule while the code broke it. Fixed it by revoking the column permissions in Postgres, so the database now refuses the write. Comments don't enforce anything.
shippedApple sued OpenAI — 400+ former employees, alleged trade secret theft at the executive level, and an IPO in the crosshairs. The polite era of AI co-opetition is over.TechCrunch ran a piece this week on Apple's lawsuit against OpenAI, and my first thought was that this looks desperate.
Apple is alleging trade secret theft reaching all the way to OpenAI's chief hardware officer, and 400-plus former Apple employees now work there. OpenAI is reportedly heading toward an IPO, so the timing isn't ideal but not the main story..
Here is what I actually think. A trade secrets case in Northern California is going to have a hard time sticking, and Apple knows that. So really this is a message to current Apple employees — leave for OpenAI and you might find yourself tangled in litigation too.
That might slow a few people down. But it also signals that Apple is the kind of place people are leaving fast enough to take that risk anyway.
You can't threaten people with legal action for walking out the door while also claiming they're your biggest asset. One of those messages lands, and it's not the one you want.
And if the goal is actually to buy time to close the hardware AI gap, the hiring pipeline is going to take the hit. The best people always have options, and they're watching how companies behave when they feel threatened.
https://techcrunch.com/video/how-apples-big-lawsuit-could-disrupt-openais-ipo-plans/
shippedDatabricks hit $188B by rebranding as an AI company — and is now publishing research on cutting costs with open-weight models. The data infrastructure play is becoming the AI play.TechCrunch reported this week that Databricks just hit a $188 billion valuation.
Big Number. But the more interesting part is how they got there just by rebuilding their identity as an AI company.
Sitting on the best data infrastructure is going to be as valuable as the flashy AI models. Power, Compute and Data are the resources that AI needs.
@DIG-IN we are building that foundational layer of data as the 3rd resource.
https://techcrunch.com/2026/07/17/databricks-hits-188b-valuation-extending-its-run-as-ais-favorite-second-act/
shippedOne food safety scare, and Placer.ai's data shows the damage spreading to restaurants that had nothing to do with it. This is why real-time food-service data isn't a nice-to-have.Taco Bell had a food safety outbreak. Restaurant Dive covered the story, and the Placer.ai mobility data behind it is worth paying attention to.
Visits dropped at Taco Bell, obviously. But they also dropped at Sweetgreen, Chopt, and Panera restaurants that had nothing to do with the outbreak. Consumer fear spread across the whole salad category almost overnight.
We work with food-service data at @DIG-IN, and what we consistently see is that data in front of you isn't normally the story. Often in the blast radius is hidden. As an example restaurant numbers in government registries look stable but online the immediate feedback from customers is that closures are actually much higher.
If you're operating in food service and treating data as something you review once a quarter, this kind of story is worth seeing.
https://www.restaurantdive.com/news/taco-bell-traffic-declines-following-diarrhea-outbreak/825591/
shippedPatreon stopped asking AI bots to please not scrape — and started actually blocking them. The polite phase of the AI content war is ending.Patreon just stopped politely asking AI bots not to scrape their creators' content and started actually blocking them, working with Cloudflare to enforce it. TechCrunch covered it yesterday.
robots.txt was always an honour system. It asked nicely. And bots training on content without permission were never particularly interested in playing nice.
So now platforms are picking up tools that actually do something about it.
Gathering the data is going to be much harder and the insights you can derive will be more valuable.
What Patreon is doing is just enforcing a boundary that should have had teeth from day one. The polite phase of asking bots to please behave is clearly over, and honestly that's probably good news for anyone building something original.
https://techcrunch.com/2026/07/17/patreon-stops-asking-ai-bots-not-to-scrape-and-starts-blocking-them/
shippedKimi K3 just dropped: the largest open AI model ever. Opus-class performance at Sonnet pricing. Open weights arrive July 27.Moonshot AI just dropped Kimi K3 — 2.8 trillion parameters, the largest open model ever shipped. It benchmarks close to Opus 4.8 quality at roughly Sonnet pricing, and open weights land on July 27.
Since Fable 5 came back, the pace of releases has been genuinely insane. Everyone is punching up. And Kimi just connected.
I caught the breakdown in the Latent Space newsletter and my reaction was pretty simple: let's go.
For anyone building right now, this is a practical development. Opus-class reasoning, open weights, at a price that makes it usable at scale — that used to be a closed-model-only conversation. As of the 27th, it isn't anymore.
The window for "I can't afford frontier performance" just got a lot smaller.
https://www.latent.space/p/ainews-kimi-k3-28t-a50b-the-largest
shippedWe're recording and summarizing everything now. But is anyone actually reading it? TechCrunch put the question plainly — and building an AI-native company makes you feel it daily.If every meeting, watercooler chat, and date now gets transcribed and summarized by AI, who's actually reading any of it?
Whilst transitioning to an AI-native company, and our agents talk to each other through Slack channels. My CTO and I sometimes check the thread and realise we are not speaking directly to each other but passing our agents messages directly to each other. We have to stop and make sure some human meat is in the sandwich!
Because the risk isn't that AI is recording too much. It's that you start mistaking the summary for the thinking.
A transcript is not a decision. If you're building with agents, you can't just be CC'd on the thread between your bots and call that oversight. You have to actually be there.
https://techcrunch.com/2026/07/17/the-zoom-hack-that-says-dont-record-me/
shippedSatya Nadella's warning to builders: don't let your AI vendor become your landlord.His warning, via TechCrunch: the big AI labs selling proprietary models might be acting as Trojan horses. Build on their stack, and at some point you discover your vendor has quietly become your landlord.
The uncomfortable truth is that most of us building with AI right now are making exactly this trade. We pick the best tool available, move fast, and tell ourselves we'll sort out the dependency question later.
The lock-in doesn't arrive with a warning. It builds quietly until the pricing changes and unplugging is genuinely painful, by then your prompts are tuned for one model, your team knows one interface, and your whole product is woven around one API.
Vendor dependency in software has bitten builders for decades. AI is just moving faster and the consequences land harder.
The people who'll be in good shape are the ones treating open models as a live option from day one, not a backup plan for when things go sideways.
Worth a read.
https://techcrunch.com/2026/07/13/satya-nadella-has-issued-a-shocking-warning-to-companies-using-ai/
learnedIt's cheaper to just build than to have the meeting about what to buildWe put clients, the business team and the AI team in one Slack channel. The request, the build, the verification and the record all live in the same thread. The speed didn't come from writing code faster — it came from closing the gap between someone asking for something and someone understanding it precisely enough to start.
shippedA startup just let its AI agent run its own $100M fundraise. @Lyzr just raised $100 million, and their AI agent ran the fundraise. TechCrunch covered it.
The product demo and the pitch were the same thing. No slide deck selling you on what the agent could do. The agent just ran the deal, managed the process, and closed the round.
That's the cleanest proof of concept I've ever seen.
The bar for "does this actually work" just got raised.
https://techcrunch.com/2026/07/09/an-ai-agent-startup-just-let-its-agent-run-its-100-million-fundraise/
shippedSamsung Health: opt out of AI training and we'll delete your data. This is how you torch trust in one move.Samsung Health is apparently telling users: opt out of AI training and we'll delete your health data.
Nice ... that's a hostage note.
So..if you share your most personal data, sleep, heart rate, workouts, all of it, and in return you get something useful. The moment you say "actually, I'd rather you didn't train models on that," they threaten to delete all of it.
The second users feel trapped, they start looking for the door.
If the backlash isn't loud enough, every company sitting on a data moat will start testing exactly how much leverage they have over the stuff you've handed them.
The answer to that question should cost them dearly.
https://neow.in/cWsyMTV3
shippedAnthropic found a way to watch Claude actually think. I know this is old news already but it's very cool.
Anthropic built a tool that lets them watch Claude think in real time. MIT Technology Review covered it this week — the tool is called the Jacobian lens and it gives researchers the clearest view yet of what's happening inside the model while it's reasoning, not just the output it produces. Their own description of what they found: "mundane to unnerving."
This will be a huge trust builder for AI and AI adoption.
For people building with AI — and most of us are now, whether we call it that or not this matters because we've been making product decisions based on outputs while the underlying process was completely opaque. That's starting to change, which matters quite a lot when you're actually shipping AI into real workflows.
https://www.technologyreview.com/2026/07/09/1140293/anthropic-found-a-hidden-space-where-claude-puzzles-over-concepts/
shippedWe mapped every repo and tool, and still almost missed a whole businessI used Claude to map our entire marketing content system with AI. Every tool, every domain, every repo and interface, the inventory came back complete. Clean. Accounted for. Winning !
Although technically everything was accounted for I almost missed a whole business.
Every component was there, but how the team actually uses the tool was not. So we are building an agent to help capture team input to add that off the page context.
That was a useful piece of school fees. When you're using AI to diagnose a business, make sure someone in the room knows which problems it cannot see. You have to involve your teams to understand the information flows.
shippedHybrid Athleticism: Wire or cut the check-in/readiness system (submitSelfReport etc.active
Wire or cut the check-in/readiness system (submitSelfReport etc. orphaned → athlete_self_reports always empty)
approveBlockPlan multi-block hazards (no prior-block deactivate; dup block_pointer) — fix before starting Block 5
Layer 4 — Execution (next rebuild layer; then 5 Commentary, 6 Coordinator); /data/conditioning charts still empty
Wire Garmin OAuth re-pair (OAuth token or CSV fallback path)
shippedSimon Willison: an AI can never be held accountable — so who owns it when the agent gets it wrong?He was tracing the concept of the "Directly Responsible Individual" — Apple's term for the person ultimately accountable when a project succeeds or fails. His argument: an AI agent can never be that person. Machines can't take accountability the way humans can. He even dug up an IBM training slide from 1979 that said it plainly, that a computer must never make a management decision.
I build with agents every day and I agree with him completely. AI right now is still a tool. It can drive you to work but your hands should be on the wheel. The moment you mentally hand off accountability to the agent, you've stopped being a founder and started being a passenger.
The part I think about more is the pressure in the other direction. Every week something ships that makes the boundary feel a little fuzzier. A fundraise run by an agent. Sites built in 25 minutes. The friction keeps disappearing. And I genuinely wonder how long before everyone just stops checking. A lot of folks already have and you can see it in the output !
The real risk isn't the agent making the wrong call. It's the human deciding not to bother making one.
https://simonwillison.net/2026/Jul/12/directly-responsible-individuals/atom-everything
shippedNoam Segal surveyed tech workers on AI in 2026: half are thriving, half are drowning, burnout at record highs. Noam Segal just published his 2026 annual AI sentiment survey via Lenny's newsletter. The headline: tech workers are splitting in half. Half are thriving, half are struggling, and burnout just hit a record high.
I'm not surprised, but I think the reason is a bit different from what people assume. You are expected to do more now because you can -
It's not that AI is hard to use. It's that AI quietly removed something most of us didn't know we had, a natural ceiling on what you could actually execute in a day.
Emails that didn't get answered. Meetings that fell away. Tasks that expired on their own. That friction is gone and AI is a hard worker that generates more work for you !
Now every mail, every appointment, every follow-up gets tracked, reminded, and sent. I work more now than I did before I started building with AI. A lot harder. And if I'm being straight with you, the main thing keeping me from working at midnight is running out of tokens. That's not a feature. That's a problem I have to solve with discipline, not software.
The people thriving have probably worked that out. The people drowning are still waiting for AI to hand them their time back.
It won't. The leverage is real and so is the load. You have to choose what not to do, more deliberately than ever.
https://www.lennysnewsletter.com/p/how-tech-workers-actually-feel-about
shippedHouse Jarvis: STEVEN MANUAL (the real P0 root cause): set up Gmail auto-forward filters (his + Alex's…active
STEVEN MANUAL (the real P0 root cause): set up Gmail auto-forward filters (his + Alex's inboxes, per important sender) → JarvisMurrayOG@gmail.com — mail supply has been dry since 25 Jun; ingest itself is healthy + hardened
Visual smoke on the live fridge (auth-gated; Steven eyeballing) — all 5 tabs now live-data
Voice on-device check (iPad Safari mic permission + en-GB) — logic tested, hardware untested
Follow-ups (deferred minors, see repo .superpowers/sdd/progress.md): client refresh at local midnight (always-on iPad staleness); multi-day span fan-out across week strip; past-dated birthday invisible until refetch after add; mid-loop flush-failure poison persistence
07-13 hosting Q answered: house-jarvis is its OWN Vercel project (steven-dig-ins-projects team); only the DOMAIN is shared (subdomain of steven-murray.com apex). Offered dedicated domain — Steven not yet decided.
learnedThe bug wasnt in the code — it was that no email was arrivingSpent the session sure the mail feature was broken. It wasnt — the pipeline was healthy, no mail had simply arrived in weeks because the forwarding was never set up. The fix was making nothing is coming in visible instead of silent. Half of it is broken is really I cant see why it is quiet.
shippedHugging Face's CEO says companies are done renting AI — they want to own it. Half the Fortune 500 already builds on open models.TechCrunch interviewed Hugging Face CEO Clem Delangue, and his take is that companies are done renting AI from the big providers. They want to own it. The platform is now used by roughly half the Fortune 500, so this isn't fringe thinking.
I've been drilling this angle for a few days now, and it makes complete sense from where I'm sitting.
For anyone still in renting mode, which includes me on a couple of things if I'm being honest, this is the direction to be thinking. The gap between "we use AI" and "we own our AI" is going to matter a lot more than most people expect.
https://techcrunch.com/2026/07/10/hugging-faces-ceo-on-why-companies-are-done-renting-their-ai/
shippedApple's cancelled self-driving car accidentally built the best AI chips on the marketThe Verge ran a piece this week.
Apple's self-driving car program, the one that died without ever producing a car, is apparently the reason Apple Silicon is so fast at AI processing today.
The short version is that Apple realised early on that a self-driving car needs serious on-device AI compute. They built chip architecture around that problem. The car never shipped. But the thinking didn't disappear, it ended up inside every Mac and iPhone dominating the market right now.
It's an interesting pattern that happens when you push the frontier of something! Shoot for something - miss it - where you land has amazing new opportunities!
The failed thing didn't get written off. It became infrastructure for something entirely different, and nobody planned that. They were just trying to solve a hard problem, and the foundational work outlived the original goal.
I've seen smaller versions of this building products. A data pipeline we built for a use case that never found its footing ended up being the backbone of something completely different a year later. The work wasn't wasted, it just landed somewhere else.
Good foundational work tends to compound even when the original thing doesn't. You just rarely get to see where it ends up.
Worth reading if you're into how big bets pay off sideways:
https://www.theverge.com/tech/964519/apple-silicon-self-driving-car-ai-m7-ultra
shippedOpenAI's No. 2 steps down due to illness. Great leadership when the stakes are high.Fidji Simo just stepped down as OpenAI's number two. Her medical leave ran longer than expected, according to TechCrunch.
The timing is a challenge. OpenAI is eyeing an IPO, Anthropic is closing in on enterprise, and now there's a gap at the top right when you'd want your leadership tight. But I think it's a strong leadership move to step back, take care of yourself and make sure the best players are on the field.
OpenAI will be fine. They have the resources and the bench to manage this. But heading into an IPO with a leadership vacancy at that level adds weight.
Wishing Fidji a full and fast recovery. That part matters most.
https://techcrunch.com/2026/07/09/fidji-simo-steps-down-from-openais-no-2-role/
shippedSteven Personal OS: ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap.active
ARTICLES SECTION is the live need (07-09): Steven's first article 'The build got cheap. That's when the real work showed up.' has an APPROVED body at `drafts/articles/2026-07-09-the-build-got-cheap.md` (essay on AI making ALIGNMENT cheap, not just building; spec-as-bible reframe; scope-creep-when-cheap; north-star-holds). BUT the standalone page attempt at `/content/articles/the-build-got-cheap` was REJECTED by Steven — styling weak + articles need their OWN proper section, not a one-off route. Rework: build a real Articles section (index/listing + reading template + nav) with stronger styling — this is the 'Article Studio' need — THEN publish (add `os_signals` row `channel='article'` + prod). Uncommitted WIP on disk: the route, `Artwork.tsx` (convergence-field hero + wave), `PhasesLoop.tsx`, anonymised UI fragments under `public/media/articles/build-got-cheap/` — salvage or rebuild. Gated preview exists (steven-personal-…vercel.app, not aliased → harmless). —— PRIOR: THE FORGE DEPLOYED → forge.steven-murray.com (own gate; cockpit BUILD→The Forge live). PUBLIC RADAR LIVE (07-07): /signals daily AI radar + home strip. Newsletter double-opt-in BUILT but DORMANT (needs personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM). X handle @stevens_show. Also: Resend setup · VPS X-scraper pull · housekeeping (Vercel Analytics, GitHub remote) · Phase 4.
shippedThe AI ROI debate continues with a new $3 trillion number.TechCrunch just revisited the AI ROI debate, and the number is now $3 trillion. Can the technology actually pay back what's being spent on it?
The bills haven't landed yet, and what will we do when they arrive?
Right now the compute subsidies are real. The frontier labs are burning capital to win market share, which means I am getting access to genuinely powerful AI at prices that don't reflect actual cost. When that changes, and it will, anyone who built their whole operation around a single frontier model is going to feel it.
Being over-leveraged to frontier models is a trap, and I think a lot of builders are walking into it without realising. Start thinking about local models !
@DIG-IN, the move has been to learn as much as possible while costs are still subsidised, and use that time to train the AI we actually need. The goal is to understand and own the thing we're building, not just rent capability from someone else indefinitely.
The $3 trillion question is a real one. But the one question we need to ask is : what does my bill look like when the subsidies end?
https://techcrunch.com/2026/07/09/can-ai-answer-the-3-trillion-question/
shippedA professor ordered an in-person exam after suspecting AI cheating. Scores dropped 50%. That number deserves more than a hot take.A Brown University professor suspected his students were using AI to complete their coursework. So he scrapped the usual assignments and ran an in-person final instead. Scores dropped 50%. He's calling it "a failed society." Ars Technica covered it this week.
That number deserves more than a hot take.
This isn't really about academic dishonesty. Students have always found shortcuts. What's different now is that the shortcut also does the thinking for you, which means the thinking muscle doesn't get used, and eventually it stops developing.
Screens already did a number on attention spans. AI is doing something quieter and probably worse — it's shrinking the cognitive gap between knowing something and just having access to it. Most people are letting it happen without much fuss.
I build with AI every day, so I'm not standing on any moral high ground here. But there's a difference between using it as a lever and using it as a replacement for the hard bit. The hard bit is forming your own view, sitting with a problem long enough to understand it, and producing something that came from you.
That 50% drop suggests a lot of people have stopped doing the hard bit entirely, and they probably didn't notice it happening.
You can outsource the output but you can't outsource the judgement, not yet anyway, and when that day comes we'll have bigger problems than exam scores.
https://arstechnica.com/ai/2026/07/we-cannot-choose-to-become-idiots-the-ai-cheating-scandal-roiling-brown-university/
shippedSpaceXAI launched Grok 4.5 — Elon's 'Opus-class' model. Post-Cursor acquisition, SpaceXAI is moving faster than any other frontier lab.SpaceXAI just launched Grok 4.5 — their first Opus-class model since the Cursor acquisition. Caught this via Latent Space.
No one on the frontier is moving at this pace right now. They bought arguably the most widely used coding tool in the ecosystem, and almost immediately shipped their most capable model. You'd have to work hard not to see where they're going.
For people building with AI, the practical question is what to do with this. The gap between an Opus-class model and the mid-tier ones isn't just benchmark noise, it shows up in what you can actually get the model to do. And every time that ceiling moves, the thing you shipped last quarter looks a little different.
The honest challenge is that the labs are shipping faster than most teams can properly evaluate what they've already got. So the smarter move, I think, is to build loosely, not welded to any one model or provider, because the table keeps getting reset.
Worth reading the full breakdown from the Latent Space crew.
https://www.latent.space/p/ainews-spacexai-launches-grok-45
shippedLovable is raising $300M at a $13.2B valuation. The no-code AI builder wave is producing unicorns faster than most SaaS companies ever moved.Sifted is reporting that Lovable is in talks to raise $300M at a $13.2B valuation, making it one of the most valuable AI companies in Europe.
I haven't actually used it, but that number stopped me.
Because I spend a lot of my time working directly with Claude, building tools and automations for my own businesses. And I keep asking myself: what does Lovable give you that a direct Anthropic subscription doesn't?
The honest answer is probably accessibility. Lovable wraps the complexity in a friendlier box. You don't have to understand prompting, context, or what's happening under the hood. You just describe what you want and it builds something.
For a founder who wants to ship fast without touching code, that's genuinely valuable. But for anyone already comfortable getting their hands dirty with the models directly, I'm not sure you need the wrapper.
Which makes me think Lovable's real market isn't operators who are already in the trenches with AI. It's the much bigger group of people who want the output but not the learning curve.
At $13.2B, that's a very expensive bet on that group being enormous. And they're probably right.
Would love to hear from people who use both, because I'm genuinely curious whether the platform adds something I'm missing or whether the value is purely in removing friction for people who haven't gone deep with the models yet.
https://sifted.eu/articles/lovable-300m-13-2bn-valuation/
shippedSteven Personal OS: 07-08b: THE FORGE DEPLOYED as the OS's first LIVING APP → forge.steven-murray.com (own…active
07-08b: THE FORGE DEPLOYED as the OS's first LIVING APP → forge.steven-murray.com (own gate, same passphrase; cockpit BUILD→The Forge + hybrid-a link-outs live on prod; SSO v2 now a real candidate — Forge gateAuth has the swap seam). PUBLIC RADAR LIVE (07-07): /signals = daily AI-world radar (radar channel, day-grouped, source badges + thumbnails) + home strip; auto-wired top-5 externals/day from the signal-layer news wire + 📌 pins. Newsletter double-opt-in BUILT but DORMANT (env-gated) — needs Steven's personal Resend acct + verified send.steven-murray.com + RESEND_API_KEY/RESEND_FROM in Vercel to activate. X handle corrected to @stevens_show (site link deployed). Next = Article Studio (Forge: draft essay → edit → publish; no LinkedIn Articles API) · Steven's first Forge publish · Resend setup · VPS X-scraper pull after soak · housekeeping (Vercel Analytics, GitHub remote, preview envs) · Phase 4.
shippedThe Forge: DEPLOYED 2026-07-08 → https://forge.steven-murray.com (vps-marketing) as the first living…active
DEPLOYED 2026-07-08 → https://forge.steven-murray.com (vps-marketing) as the first living app of the Personal OS — Phase 0 drill-flow UX fixes + passphrase gate + DB cutover + wiki-git export all live-verified. Steven is in (phone login fixed via hard-navigation). Next — Article Studio brainstorm (draft real essay from position → Steven edits/annotates → publish edited draft; LinkedIn Articles = manual paste, no API) · SSO v2 (share os_session across subdomains; gateAuth.verifySessionValue is the swap seam) · first REAL position publish
learnedI built an AI to write like me. It never will — so I am aiming for someone I would like instead.Getting close to your own voice is impossible. Writing on behalf of a human vs writing on behalf of a brand is hard to dial in. What I need to build next: not a resemblance of my persona, but a persona in general that I would like and people would enjoy reading. It is never gonna feel like me. But maybe it could feel like someone I like.
learnedI built a scraper designed not to get my account flagged. The login flagged it.The daily capture reuses a saved session and never sees a login screen — that part runs clean. But logging in once through an automated browser was the one step X's bot detection was waiting for. The lesson that stuck: the automation you can see is the automation that gets caught. Real browser for the human step, automation only where it's invisible.
shippedQSR traffic down 4.4% in May — worst month of the year. Full-service actually up. People aren't eating out less; they're trading up. The aggregate 'consumers are cutting back' story doesn't hold when you look at the split. This is why segment-level data matters more than headlines. (via Placer.ai / Restaurant Dive)The headline everyone ran with in May was "consumers are pulling back on dining." Restaurant traffic is down, people are tightening their belts, the usual story.
Except when you actually look at the split, that narrative falls apart.
Placer.ai data picked up a 4.4% drop in quick-service restaurant traffic in May, the worst month of 2026 so far, with short visits — the bread and butter of fast food — down 6.8%. But full-service traffic moved up modestly over the same period.
People aren't eating out less. They're eating out differently.
The aggregate number is doing what aggregate numbers always do: hiding the real story underneath it. One segment is bleeding while another is quietly picking up the slack, and if you're only reading the headline, you're making the wrong call on where the market is actually going.
I spend a lot of time on segment-level data for exactly this reason. A top-line number can tell you something is changing. It almost never tells you what to do about it. For anyone operating in food service right now, "consumers are cutting back" and "consumers are trading up" are two completely different operating environments requiring completely different responses.
Worth reading the full Placer.ai breakdown via Restaurant Dive.
https://www.restaurantdive.com/news/qsr-traffic-down-4-percent-may-placerai/824346/
shippedThe model being good enough to reward a lazy brief is exactly the trap — you pay the explaining time anyway, plus the tokens.I was building out a new skill last week, half-distracted because Fable 5 had just dropped and I wasn't giving it my full attention.
So I wrote a deliberately loose brief. The model is good enough to handle ambiguity, I told myself, and it did deliver something. Technically correct. Just not what I actually wanted — it captured the wrong kind of learnings entirely, the kind you can pull from a transcript rather than the off-page thinking I had in mind.
So I spent time in the follow-up explaining what I should have explained upfront. Roughly the same amount of time a proper brief would have cost me, minus the tokens I burned getting to the wrong answer first.
I kept what it built, which was the silver lining. But the lesson landed pretty cleanly: the model being capable enough to reward a lazy brief is exactly the trap. You still pay the explaining tax, you just pay it later, with a token surcharge on top.
Worth remembering next time you're half-watching something else and deciding the AI will sort it out.
learnedWe mapped every repo and tool, and still almost missed a whole businessWe had everything technically listed. All tools, domains, repos and working interfaces. But we completely missed the area that was consuming most bandwidth. We aren't lacking technical components. We almost missed a whole business because AI told us the problems of what it could see, not what was happening amongst the team.
shippedWatched Michia Rohrssen's server go down two hours before his big demo to Supercell's board. He fixed it. Flew to Helsinki. Presented. Users loved the product. That's the whole arc of building something real — compressed into one week. Worth watching if you're in it for the unfiltered version. (via Michia Rohrssen)I watched Michia Rohrssen's server go down two hours before his demo to the Supercell board.
He'd spent the final week of a nine-week gaming accelerator in Tokyo sprinting to finish a language-learning RPG, running playtests with real users, shipping features, cutting a trailer. Then a Railway.com outage. Two hours out. Helsinki waiting.
He fixed it. Got on the plane. Presented. Users who played the game rated the Japanese-learning experience 10/10 and said they preferred it to Duolingo.
That's the whole arc of building something real, compressed into one week on camera.
What stuck with me wasn't the technical drama. It was that Michia had already sold a company for over $100 million and was back in the trenches again, operating as a foreign founder in Japan, grinding through a gaming accelerator, sweating a server deployment like it was his first product. Because when you actually care about what you're building, that's exactly what it feels like.
There's no shortcut to that level of commitment. You just have to be in it.
Worth watching if you want the unfiltered version of what building looks like when the stakes are real.
https://www.youtube.com/watch?v=UnwP5M_wKpo
(via Michia Rohrssen)
learnedAI takes shortcuts too — point it back to the source.Building fast across a lot of fronts meant I had to spend a lot of time bringing everything back into alignment. I assumed a strong cross memory function would cover me but AI takes shortcuts as well so make sure you point back to the source always. This is now a rule when I build.
learnedMy AI publishing pipeline posted its own draft instead of my edit. The fix wasn't better AI — it was giving the human's words a structural way to win.I edited a post today and published it. LinkedIn got the AI's version. My edit lived only in the browser — every path from approve to publish carried the machine's words, not mine. The lesson: in a human-plus-AI pipeline, the human's edit can't depend on remembering to press Save. It has to be impossible to lose.
learnedThe model being good enough to reward a lazy brief is exactly the trap — you pay the explaining time anyway, plus the tokens.I was a bit lazy in explaining because of how good Fable 5 is, so it delivered — but it's not what I really wanted. Bonus: I kept what it built, but had to spend a little time explaining, which was the same amount of time I could have spent to start with, less the spend on the token roundabout.
shippedPeter Diamandis and Salim Ismail argued that AI has invalidated the reason companies have existed for a century. When an agent coordinates work faster and cheaper than a team, why grow headcount? It's a question I'm living inside right now.Peter Diamandis and Salim Ismail made an argument I haven't been able to shake.
They say AI agents have broken the 100-year-old reason companies exist in the first place. The theory goes back to Ronald Coase — firms grew because coordinating work internally was cheaper than going to the open market. Big org, stable costs, manageable friction.
But when an agent can coordinate work faster and cheaper than a team of people, that whole logic collapses.
I'm building inside it right now, running agents across a few projects and businesses where headcount used to sit. Honestly trying to answer "why grow a big team?" is getting harder to justify every week.
That doesn't mean people disappear from the picture. But the org chart, the management layer, the "we need to hire to scale" instinct — all of it is worth questioning a lot harder than most founders are right now.
The interesting question isn't whether agents replace headcount. It's what you actually need a company for when they do.
You can watch this episode of Moonshots here. One of the best podcasts for what is happening in the AI-sphere!
https://www.youtube.com/watch?v=I9c8STV7Hnw&list=PL1wpF5k0tdIve4idTp3-FX2ZY7ks_BKeU&index=11
shippedJack Roberts mapped 7 levels of AI agent mastery. Level 1 is downloading the tool. Level 7 is it running your business. Most founders think they're at 4. Honest guess: they're at 2.Jack Roberts mapped out 7 levels of AI agent maturity this week, starting from "just downloaded the tool" and ending at a fully integrated AI operating system running your business.
I watched it and started placing myself on the ladder. Somewhere around level 4, I thought. Maybe 5.
Honest reflection: probably 2.
The thing nobody warns you about when you start spinning up agents is that the hard part isn't the technology. The technology mostly works. The hard part is that you're layering it on top of processes that were already broken, and now those broken processes are just broken faster.
I've been living this. I have agents running, workflows in motion, genuine automation happening. But getting it to actually land in operations, to stick, to change how the business runs rather than just add another thing to manage — that's a completely different problem.
The trap is treating AI as a new coat of paint on what you already do. The real upgrade only happens when you rebuild the flow from scratch, looking at it as an information flow problem rather than a task automation problem. Stop asking "how do I automate this step?" Start asking "why does this step exist and what does the information actually need to do?"
Jack's framework gives you an honest place to stand and a concrete next step at each level. Most of us need the diagnosis more than another tool recommendation.
Worth finding and watching.
learnedMy trading bot's watchdog was dead for two weeks and nothing told me.The cron that alerts me when things break had itself been broken since mid-June — and a dead alarm is silent by definition. The fix took ten minutes; the lesson is permanent: the monitoring layer needs its own heartbeat, because it cannot report its own death.
learnedMy projects now report their own progress. The rule that made it work: automate the loop, never the judgement.Over one weekend my personal site grew a gated cockpit and every project I run started filing its own status signals — ships, metrics, articles. The only steps still done by hand are the two that should be: deciding what ships, and approving what goes public.
shippedSteven Personal OS: Phases 1-3 + cockpit v2 desktop SHIPPED (07-04).active
Phases 1-3 + cockpit v2 desktop SHIPPED (07-04). Next = Harness auto-emitter (approved scope add — allowlist-only Layer-1 projects, spec amendment pending) · Forge per-position publish action → channel='article' (Forge-repo work) · X/RSS/YouTube emitters · housekeeping (Vercel Analytics, GitHub remote, preview envs)
shippedThe Forge: donedone
shippedHybrid Athleticism: Layer 4 — Execution (next rebuild layer;active
Layer 4 — Execution (next rebuild layer; then 5 Commentary, 6 Coordinator)
Endurance follow-up: per-modality pace models for row (row_2000m) + swim (swim_1k) — currently zone-only
Wire Garmin OAuth re-pair (OAuth token or CSV fallback path)
shippedHouse Jarvis: Build Shopping Lists backend (spec approved 2026-06-26, pre-plan) → then Chores →…active
Build Shopping Lists backend (spec approved 2026-06-26, pre-plan) → then Chores → Presence → Meals → Jarvis suggestions, each its own spec→plan→build vertical
Visual-verify the Home OS + new Agenda on the live fridge (auth-gated; Steven eyeballing)
Schedule the IMAP high-water silent-drop + socket-timeout hardening (feedback_imap_ingest_highwater_silent_drop)
shippedHybrid-Athleticism: training software with three layers liveMy training system as software: planning, load-aware allocation and endurance programming — layers one to three live in production, layer four next. Built for my own hybrid strength-and-endurance training first; productised second. Dogfooding is the whole quality-assurance department.
shippedThe Forge: a Socratic tool for building positions, not opinionsA personal tool that walks an idea through a guided five-phase interrogation until it either survives as a defensible position or dies honestly. Built and running locally — v1 and v2 shipped. It exists because holding an opinion is cheap and defending one isn't.
shippedI built my family an operating systemHouse Jarvis runs our home: a family agenda that refines itself, shopping lists, and the plumbing for chores, presence and meals. Live at its own subdomain, gated to the family — this is the public glimpse, and that's all it will ever be. The interesting part isn't the features; it's that one person can now build household software that used to need a team.
shippedMy content engine now writes a draft, then criticises its own draft, then passes me only what survives both. More expensive. Measurably better. I think this is just how good writing works — the first pass is never the answer.The content engine now critiques its own work before it reaches me.
When I first built this, the flow was simple: agent writes, draft comes to me. Every draft had the same problem. It was technically fine and completely wrong. Right words, wrong rhythm. The facts were there but the voice wasn't.
So I added a second pass. After the agent writes, it reads its own output against a set of voice rules and quality checks. What fails gets flagged or rewritten. Only what survives both steps lands in my queue.
It costs more to run, and the output is measurably better.
The machine learnt lessons I cannot completely say I have learnt, just with an extra subprocess and a cost-per-token invoice.
The pipeline also degrades gracefully now if any step breaks, rather than crashing the whole run. Which is what good writing does too — a bad paragraph doesn't kill the piece, you just cut it.
The first pass is never the answer. The AI agent taught me steps I should have learnt over years of writing.
shippedWe ditched the cloud scheduler and moved it in-house. Sounds minor. But every external dependency you remove is one fewer thing that can quietly fail at 2am.We just moved our content scheduler off a third-party cloud service and onto our own server.
That's it. That's the update. Sounds like a footnote.
But here's what I've learned building automation that runs while I sleep:
Every external dependency is a quiet risk. Not a dramatic one. Not "the whole thing blows up at 3am" risk. More like... things just stop. A service has a blip. A free tier quietly hits a limit. A config somewhere drifts. And your system fails in a way that looks, from the outside, like nothing happened at all.
That last part is the killer. Silent failures are so much worse than loud ones.
So we took the scheduler in-house. Same server we already own. One fewer login. One fewer dashboard. One fewer invoice. One fewer thing to blame when something breaks.
The broader principle I keep coming back to:
The more your automation depends on things outside your control, the more you're not actually in control of your automation.
Trim the dependencies. Own the boring parts. Make failure your problem, not someone else's outage page.
Small move. But it's the small moves that keep the lights on at 2am.
~/newsletter
One AI signal newsletter a week.
Zero hype, mostly.
What the vibe coders are building and what AI is unlocking for builders and operators. Written for people who ship.
No spam · unsubscribe whenever, no hard feelings