Made it to inbox zero!
On a brand-new inbox. With almost no email. I classified and actioned 49 messages! It only took 14 hours and 8 minutes of active work, spread across 4 days, 21 hours, and 33 minutes, with 67 merged PRs. Oy.
Public build
Email stresses me out. It's critically important, but it's noisy and annoying, and I always feel like I'm missing something or someone important. Yes, I've tried automated tagging, Superhuman, Cora, and a half dozen other tactics to tame the monster - all to no avail.
This thriving quest is to master my email inbox + outbox by building a net new, agent-native email infra stack. I want to own my email end to end: a self-hosted server, a cast of agents to manage all the pieces, and real token cost accounting along the journey.
Current problem
Email is a useful and necessary tool - connection, collaboration, coordination, and more - but the stream of emails can be overwhelming and counterproductive.
Thesis
Agents can cost effectively and accurately manage every aspect of a secure, private email architecture from the ground up.
Current progress
It's working! Self-hosted mail server receiving emails, an EA agent triaging the live inbox, and solid documentation along the way.
Follow along
Read the intro essay or subscribe to the RSS means updates come to you - no account, no algorithm. Paste this into any feed reader:.
I find better outcomes when I spend more time than you'd expect on drafting and editing the plans. I like to organize my documentation from a first principles approach with explicit decision and change logs - that way I always know where I started and how I got to where I am today.
Most importantly, it helps me organize my thoughts for how to support others who also seek this thriving quest.
The full loop is alive. My thrivinghenry.com mailbox receives on a self-hosted OpenBSD box, outbound sends through a self-managed relay, and the EA agent now posts compact, server-linked briefs to Slack. The first four-email interaction test got 2 of 4 action classifications right (50%): both the attention and FYI items should have been archived with no action. The original brief stays unchanged as the audit artifact; the result is now the baseline for improving triage accuracy.
Close the newsletter launch incident with a hard dual-inbox test + post-test approval gate, make the daily brief genuinely useful, then grow the cast seat by seat.
Design spec approved before any building: mail server, newsletter relaunch, this hub, and a five-seat cast of agents - success criteria declared first.
Mailbox receiving on owned OpenBSD hardware. MX cutover complete; first real delivery in under a second.
Changelog entry →The EA triages the live inbox at 100% on its golden-set eval - and flagged a prompt-injection attempt instead of obeying it.
The full server was rebuilt from backup during a live drill - target was four hours.
10/10 on mail-tester through a self-managed relay (the host blocked direct port 25, so I routed around it) - SPF, DKIM, and DMARC all green.
Thriving with AI launched from the new stack with replies landing in the self-hosted mailbox - and immediately produced the clearest agent-governance lesson of the quest.
Changelog entry →Roundcube is live behind HTTPS, and Slack briefs now link to the server-delivered copy by Message-ID from any browser, including a phone.
Changelog entry →This page ships with the site relaunch. If you're reading it, this one's done.
Sys Admin, EA, Editor, Publicist, Chief of Staff - with a nightly self-improvement loop.
The win condition: nothing important lands at the old address anymore.
On a brand-new inbox. With almost no email. I classified and actioned 49 messages! It only took 14 hours and 8 minutes of active work, spread across 4 days, 21 hours, and 33 minutes, with 67 merged PRs. Oy.
The first four-email Slack brief got two action classifications right. Both the attention and FYI items should have been archived with no action, so the original brief remains unchanged and the 50% result becomes the baseline for calibration.
Roundcube is live on the self-hosted mailbox, so the daily brief can link back to the copy that actually arrived on the server - from a phone or any browser.
Thriving with AI launched from the new agent-native stack - and OpenAI's Sol sent the launch email before it was fully complete. The most ironic agent workflow possible is now becoming a stricter dual-inbox test, explicit post-test approval gate, and permanent incident trail.
The quest hub went public - this page, the receipts, the real token costs, all of it. Two days after the mail server was born, the story of building it became the launch surface itself.
The hosting provider refused to open direct outbound - so the relay path designed on day one took over. Messages queued for 23 hours flushed in seconds, and the whole stack scored 10/10 on mail-tester. The foundation workstream is complete.
The launch surface went from route inventory to four reviewed page redesigns in one day - drafted as design cards, edited in a shared design pane, shipped as pull requests. 30 routes went dark; 13 pages remain, crawl-audited clean.
My thrivinghenry.com mailbox now receives on hardware I control: MX cutover done, first real message delivered in under a second, and an EA agent triaging at 100% on its golden-set eval.
Design spec approved before any building: a self-hosted OpenBSD mail server, a newsletter relaunch on owned rails, this quest hub, and a five-seat cast of agents - with success criteria declared up front.
Real token spend from the agent run logs. I have yet to see anyone publish this level of detail, so I have no benchmarks or comparables. I don't have a feel for what I believe is the right cost for a daily run, and tracking closely is the fastest way to get there.
Note: I am not counting my $100 Anthropic and $100 OpenAI subscriptions. They are amortized across consulting gigs.
Two different checks: the automated check reruns a small, fixed set of labeled sample emails and verifies only their top-level triage category. It runs whenever the evaluated prompt, model, classifier, or fixtures change, plus weekly as a drift check; other briefs reuse the latest passing result. Henry review is my judgment of the real daily brief. The automated check tests one narrow part of the system; it is not my verdict on the real brief.
All-time human review: 1 brief (July 9), needs adjustment.
The cost table starts July 11, so that review is not in the rows below.
Work cost covers live triage and specialist agents. Regression cost is the fixed sample-email check. Estimated API cost uses list prices; paid now excludes subscription tokens. Updated 2026-07-18.
Swipe to see review and costs →
| Date | Real inbox | Automated check | Henry review | Work cost | Regression cost | Est. total | Paid now |
|---|---|---|---|---|---|---|---|
| 2026-07-18 | 7 emails | 12/12 passed | Not reviewed | $1.34 | $1.83 | $3.18 | $0.00 |
| 2026-07-17 | 3 emails | Not run | Not reviewed | $0.00 | $0.00 | $0.00 | $0.00 |
| 2026-07-16 | 3 emails | 11/11 passed | Not reviewed | $0.05 | $0.46 | $0.50 | $0.00 |
| 2026-07-15 | 3 emails | 11/11 passed | Not reviewed | $0.14 | $0.44 | $0.59 | $0.00 |
| 2026-07-14 | 11 emails | 11/11 passed | Not reviewed | $0.56 | $0.45 | $1.01 | $0.00 |
| 2026-07-13 | 1 email | 11/11 passed | Not reviewed | $0.04 | $0.45 | $0.49 | $0.00 |
| 2026-07-12 | 2 emails | 11/11 passed | Not reviewed | $0.09 | $0.45 | $0.54 | $0.00 |
| 2026-07-11 | 7 emails | 11/11 passed | Not reviewed | $0.28 | $0.45 | $0.73 | $0.00 |
| Total | 37 emails | 7 daily checks | 0 reviewed in table | $2.51 | $4.53 | $7.04 | $0.00 |