TWA-005September 18, 2026

Three frontiers and a lake

Hey hey,

Big, big news: my @Work page is live!! This has been a labor of love for the last few weeks, and more on that in a second.

First, the last two weeks. Big leaps at the keyboard, with solid family time outdoors in between. We visited a family friend near the South Yuba River, then went camping at Sardine Lake. Kiddo is picking up many new skills and it was extra fun to see him tromping around a campsite. Lots of good nature in these parts, and I’m so grateful for the time to rest my eyes and my spirit in the natural world.

Thriving with AI @Work is live

Hype aside, AI is reshaping any work that touches a keyboard. Thriving with AI @Work is my hub for sharing the patterns I see work well. It’s organized as three frontiers, and one is more polished than the rest.

Start with the second one. AI Transformation is the tightest of the three by a wide margin. The premise: AI transformation starts with what a business could become when its people can do more better. From there it walks through six principles, from “Let the business outcome lead” to “Lead the transition.” Two companion pieces go deeper: a field guide for getting from a promising idea to a practice that lasts, and an essay, More output is only the beginning. The sources are primary and linked in the notes, so you can always read deeper.

The other two are up, and they are sketches:

Both show only the most rudimentary outline of what I’m excited to share on those fronts. I’ve been deep in it and wanted to share my progress so far.

In motion

An epic knowledge base build

So … anyone else notice that tokens were fast and cheap over Burning Man weekend? When half the Bay Area’s AI crowd is out on the playa, let them tokens rip! I made the most of it and built a knowledge base where everything I read, watch, and listen to compounds instead of evaporating.

I was triggered by a very annoying trend I’ve noticed - I cited a roundup article on my site, and when I traced its numbers back to its citations, several were misquoted and a few were simply made up. Sloppy mcSlop barf. I want that to be impossible by design on my site, without relying on me being vigilant.

So every source gets snapshotted with its date and a hash, in case it changes or vanishes. Every idea I pull out is its own small file with the exact quote and where it lives in the snapshot. Data points only count when they come from primary research or a first-hand account. People, companies, and topics are three views over the same evidence, so nothing gets copied three times. It’s all flat Markdown files. No vector database, no graph database, nothing hiding the data from me. A checker fails the build when a quote isn’t actually in its source.

First merge on September 7 held 34 sources and 33 extracted ideas. Today the checker reads 6,990 sources, 4,709 extracted ideas, and nearly 500 thinker profiles at varying depth across 14 functional areas. All seeded with the full back catalogs of a few writers and publications I’ve grown to trust over time. It also marks what I personally have read versus what only my agents have, which keeps me honest about whose opinion I’m repeating. And it bubbles up things that are specifically relevant for me to read based on the challenges and explorations I have right now.

I’m absurdly stoked on how this turned out, and it’s already been immensely helpful in a few tactical domains. I’m very excited to see how this continues to grow and compound in my world.

A family reunion, run as a workstream

We’re planning a big family gathering over the holidays: about 20 people confirmed, maybe 30, ages 1 to 80+, plus two dogs. Finding a place that sleeps everyone, has one room where we can all sit down to dinner, is gentle on the elders, has entertainment for various ages, and welcomes a very large dog is the kind of search that usually eats a month of evenings and a sprawling group text.

So on a Sunday night I spun it up as a multi-agent workstream instead. About 90 minutes in, I had a project folder, written hard and soft criteria, a first round of availability checks run on each property’s own booking page, and a weighted scorecard in a shared Google Sheet the family can poke at. I ran two frontier agents on the same problem in parallel, then had each one review the other’s work and folded the best of both together.

Where it stands, five days from the start:

  • 75+ options evaluated: 27 scored against our criteria, about 50 ruled out, each with the reason written down. That ruled-out table is reusable for events after this one.
  • 20 distinct plans ranked into one action queue.
  • 21 inquiry emails drafted overnight and waiting in my drafts for review. Agents draft and track the threads so every question gets fully answered in a timely manner, and I hit send
  • 4 strong options with pros and cons at various price points
  • 1 top option modeled out with per family costs

Nothing is booked yet, so need to cross the finish line, but the amount of progress in a simple week with 10-15 minute stretches between running after my kiddo and work meetings is so wild.

Taking in

Been reading about and thinking about work stuffs … can you tell?

Notes on Building a Code Reviewer (opens in a new tab) by Dana Dzik at Coval. Really interesting and unexpected pattern where a roster of 4 “dumber” model runs caught more defects than 1 run of the strong model (19 vs 16 on their test set), at roughly a third of the cost. Also, cool to see the coding equivalent of what I’m seeing with my kb and writing - forcing agents to cite their sources, and then check the citations are real, cuts hallucinations way down (98% of their cited claims pointed at real code). Forcing citations is the most useful pattern I know for cutting hallucinations.

Do Automated Evals Work? (opens in a new tab) by Antaripa Saha and Hamel Husain. They tested tools that promise to analyze an agent’s failures automatically, and every one missed the same kind of miss: the trace looks correct, but the agent falls short of the product’s real goal. I expect smarter, more expensive agents could do better, though honestly the pattern above that Dana is building with Sofia is definitely the answer. Plus, having a kb with known best practices and things to look for would really supercharge this analysis (also something we’re building for). Their sharpest line is that no system did the one thing that would have helped most: interview them. That’s one of the meta patterns that I’m building into the Human-AI Collaboration frontier.

Why eval startups fail (opens in a new tab) by Thomas I. Liao. From last year, and I work at an evals company, so of course I read it with a pen in hand. His case: good eval talent drifts to better-paid parts of the stack, and the customers who need evals can mostly build their own. Fair on talent. I’m less sold on the customer point. In voice, the group of teams building on models who can’t cheaply evaluate them is a lot bigger than he suggests, and even teams who could build it face a real opportunity cost. He closes with the distinction that matters most to me: selling evals and selling evals tooling have very different economics. Bonus points for the Goodhart’s Law reminder: when a measure becomes a target, it stops being a good measure.


This coming week I’m keen to make some progress on my writing, as well as some of my personal Thriving Quests. And a few more things I want to fold into the kb before I’m done with this first wave. So much fun … mo’ tokens pweaze :)

Until next time,

~h

Join the exploration

Thriving with AI newsletter

A roughly weekly newsletter where I share what I'm learning and building at the frontier of working with AI @Work and @Home.

Join now for early access to what launches in October and November.

Free, with a 100% money-back guarantee within 90 days.