Build the work people can actually use
A convincing demonstration answers a small question: can this work once? I want to answer the next one: can the people responsible for the business use it, trust it appropriately, and improve it when circumstances change?
The useful entry point depends on the constraint. I might start with a dubious outcome, a capable tool that people don’t use, or a team whose experiments never become dependable work. These choices overlap; they aren’t a mandatory sequence. I return to them as the evidence changes.
Find the constraint worth changing
If this is you: the requests are for features, the results are supposed to be a business outcome nobody has named, and every conversation keeps turning into a conversation about tools.
I start with a business outcome: growing revenue, reducing costs, accelerating innovation, or reducing a material risk. Then I follow the work backward. Where does a customer wait? Where does a good idea stall? Which decision takes longer because the person making it can’t find what matters?
A workflow observation often tells me more than a list of requested features. I ask someone to walk through a recent example, including the workaround they barely notice anymore. I want to see the handoffs, the source material, and what happens after the output leaves their desk.
Felt pain is one starting point. Newly feasible work is another: an experiment, service, or analysis that used to cost too much to attempt. I want an intended beneficiary and a decision worth making.
A useful signal: if the requested feature has no clear user or downstream decision, I’d observe one recent piece of work before building. If that reveals a cross-team handoff as the constraint, I’d widen discovery to include the people on both sides.
I keep discovery proportionate to the uncertainty. One observed workflow may be enough for a bounded test; a cross-functional change needs the handoffs, authority, and downstream effects in view. I widen the inquiry when the apparent constraint moves, rather than mapping the whole company by default.
I also inventory what already works. An existing tool may need a clearer instruction, better information, or a dependable handoff. Rebuilding it would spend time, discard local knowledge, and give its owner another thing to learn. New infrastructure has to earn its place.
I look for recurring pain, available information, manageable consequences, a way to evaluate the result, and someone who will own it. That helps me choose something important enough to matter and bounded enough to learn from. A narrow first experiment can serve a large ambition.Natalia Quintero’s executive guide (opens in a new tab) draws on Every’s experience of abandoning an overly broad project-management automation. Her screen includes frequency, pain, data, risk, ownership, evaluation, and maintenance. I like that it makes the work after the demo part of the initial choice. These are prompts for a discussion, not validated scoring weights; a tidy score can conceal a disagreement about what matters.
I like to write the expected path to value in plain language before choosing the intervention. A team spends less time preparing for customer conversations, uses the capacity for better conversations, learns what matters sooner, and earns more business. Each connection in that chain is a hypothesis to inspect. A faster task can leave the commercial constraint untouched.In GTM in 2025 (opens in a new tab), sales leader and investor Sam Blond challenges the automatic inference that time saved becomes revenue. I agree with that qualification. His practitioner argument helps frame the question here: which commercial constraint does the saved time actually change? If the team already has enough preparation time but struggles to reach the right customers, more prepared material may change very little.
Much of the craft is dissecting the full workflow before and after a selected use case. I look at how decisions get made: by whom, in what room, and based on what information. I consider how that information gets collated, from what sources, with what logic, and with what editorializing by the analyst. I dig into where individuals wait, hand off work, and hit roadblocks that prevent useful work from finishing.
Make excellence available to the work
If this is you: the same corrections come back every week, and the two people who know what good looks like are explaining it one colleague at a time.
An experienced colleague knows which detail makes a customer response useful, when a recommendation needs more evidence, and which exception deserves a conversation. A model doesn’t arrive with that organization-specific judgment. I make as much of it explicit as the work requires.
I start with expensive repeated corrections, rather than asking a team to document everything first. Instructions need to be ready to act on at the point of work: which context applies, what to do, what to check, and when to ask for help.
I ask for examples of excellent work and plausible work that missed the point. The explanation matters: why was one useful, what was missing from the other, and who can authorize an exception? Those distinctions become shared guidance for people and machines.At Coval, our customer-facing materials used to look different depending on who made them. Our lead designer created a shared design system, and the work became more unified and polished overnight! Across the whole team! So cool. That experience is a small example of the larger idea: make the conventions behind good work available to everyone. The same question is useful in finance, sales, and engineering: what does an experienced colleague know that the rest of the system keeps having to rediscover?
I treat this guidance as an operating responsibility: name its owner, make the current version easy to find, and retire contradictions.In her conversation with Lenny Rachitsky (opens in a new tab), Jessica Fain asks how we would onboard a hundred new colleagues who do not know our product philosophy. Around 1:18:47, she turns that question into practical work: articulate what matters, how decisions get made, and where people need to catch mistakes. Her onboarding analogy resonates with my view of agents as brilliant newcomers with no organizational context. Named ownership and retiring contradictions are my extension.
I distinguish a lasting business standard from a temporary workaround for a particular model. Both may be useful today; they need different maintenance.
More documentation can make context selection harder. The system needs to know which team’s conventions apply and whether the instructions are still current. I want a freshness check and a way to settle conflicting instructions.Gal Bakal’s Knowledge Activation paper, sections 4.1 and 8.7 (opens in a new tab), describes a very recognizable failure: overlapping team conventions make plausible guidance easy to select incorrectly. Installed copies can also fall behind the maintained version. The lesson I take is that documentation quality includes selection and delivery. Publishing a better instruction is only part of the repair if people and agents keep using an older one.
Codified practice can preserve a bad habit just as efficiently as a good one. I distinguish the instruction to follow a standard from the assignment to challenge it, and make proposed changes reviewable.Bakal discusses this tension in section 9.4 (opens in a new tab): reusing established knowledge can come at the expense of exploration. A convention has a history, and sometimes that history is an old constraint that no longer applies. My rule of thumb is to give improvement its own explicit assignment, so dependable execution and questioning the method both have a place.
Give judgment a clear job
If this is you: nobody is sure when to trust the output, who can override it, or who owns the decision.
I use AI to gather and synthesize information, develop alternatives, draft, code, and advance separable work in parallel. That gives me more room for judgment and creativity. In an organization, I want the same design to amplify people’s domain expertise and their relationships with customers and colleagues.
I choose among three useful arrangements by task consequence:The Enterprise AI Playbook (opens in a new tab) is a Stanford Digital Economy Lab study by Elisa Pereira, Alvin Wang Graylin, and Erik Brynjolfsson, drawing on interviews with executives and project leaders about 51 successful AI deployments across 41 organizations. Pereira was an MSx candidate at Stanford Graduate School of Business (GSB). On page 30, the authors distinguish collaboration, approval, and escalation. The labels below are my practical rendering. I find the distinction useful because one organization can need all three, even within a single workflow. These selected success cases do not establish that greater autonomy causes better results; task selection and consequences matter.
Where does the person enter the work?
In every arrangement, a person sets the outcome, context, and boundaries.
Work together
- Person frames the task
- AI explores and drafts
- Person judges and redirects
Explore, challenge, and revise together.
Work through an ambiguous customer problem.Approve then act
- AI prepares
- Person approves
- Action
A person decides before a consequential step.
Review a proposal before sending it.Act within limits
- AI acts within limits
- Exception?
- If yes: pause for a person
Otherwise, routine work continues within the agreed limits.
Update approved fields; escalate a conflicting source.The first arrangement echoes the “human sandwich” described by Dan Shipper in After Automation, crediting Kieran Klaassen: people frame the work and use judgment to shape what follows.
Observable outcomes and reversible mistakes can support more delegation. Hard-to-detect errors or costly recovery call for tighter limits. These are operating choices, not stages everyone should progress through.
I specify permissions in the operating system and tools, not just in prose. A written instruction to seek approval is not an enforced permission boundary. Checks should make their coverage clear; human accountability remains.Bakal’s validator proposal in section 5.5 (opens in a new tab) makes particular properties testable. That is useful, but it is easy to overread the reassurance of a green check. For example, a correct format says nothing by itself about permission to send a message. That example is my distinction between validating an output and authorizing an action; passing selected checks does not cover every important risk.
These boundaries can change as capabilities improve. I want evidence for revising them, with a named person still accountable for the result.
Evidence needs special care. An earlier human opinion can return through a summary looking like fresh corroboration. I separate primary facts, prior judgments, and the model’s new conclusions so the reviewer can tell whether anything independent has actually been learned.
Historical examples need the same discipline. When testing a past decision, I restrict the inputs to information available at the time and keep the answer and later summaries out of reach. Agreement with the old decision is one observation; factual accuracy, sound reasoning, and usefulness deserve separate assessment. People may have been wrong, too.
Make the experiment safe to learn from
If this is you: the same few people are carrying every pilot, delivery is slipping, and nobody wants to be the one to say a bet isn’t working.
When the portfolio is tiring people out: if the same champions support several pilots while delivery slips, I’d review those bets together. Compare each with the business priority and evidence so far; pause or stop weak bets before adding another. A smaller next experiment should have an owner, a review date, and an explicit capacity decision.
Before building, I write down the hypothesis and the decision the experiment should inform. I establish a baseline, a quality bar, and a review date appropriate to the question. A quick operational improvement and an exploratory innovation bet need different horizons.
I define what the system can read and what it can change, including the point where a person must approve an external action. I want mistakes to be recoverable within the experiment’s boundaries. The level of autonomy follows the consequences of the work.
Then I test representative cases, awkward exceptions, and known misses. I keep examples used to tune the workflow separate from examples reserved for evaluation. Repeatedly improving against the same cases can teach the system the test without teaching me much about its broader usefulness.
Leaders must make capacity for the test: reprioritize existing work, add capacity, or agree to a bounded stretch. I don’t prescribe a universal allocation or promise learning will pay back on a fixed schedule. Review and maintenance keep taking time after the first useful result. Then the nature of work itself changes.
A useful experiment can end in revision or a decision to stop. I want people to describe failures candidly, including their own choices. That requires leaders to welcome the information and respond constructively. Permission to learn has a purpose: making the next business decision better. I celebrate candid learning from misses while still requiring useful successes across the effort. A review should explicitly continue, change, or stop the experiment.Chapter 4 of the Enterprise AI Playbook (opens in a new tab) describes protected learning and continuity of sponsorship in its selected cases. I take those accounts as a reason to ask what happens after an experiment disappoints: is there still time, support, and permission to tell the truth? These are largely leader accounts. They do not independently establish how employees felt or prove a formula for psychological safety.
Raise the floor, then practice the craft
If this is you: one person built something remarkable, and the next person can’t use it without them in the room.
I want everyone to share a starting vocabulary and a basic experience of useful AI work. That learning floor includes making an attempt, checking the result, redirecting it, and recognizing when to ask for help. A common introduction creates a place to begin together.
I separate two learning streams. One spreads workflows that help today. The other grows people who can exercise judgment tomorrow, including when a first draft or a routine analysis is no longer theirs to produce.
Spread the workflow. Grow the judgment.
Useful work, shared today
- Find a workflow that helps.
- Show a colleague how it works.
- Package the examples and instructions.
- Support repeat use and improve it.
More people can do useful work now.
Expertise, earned over time
- Give someone a real decision.
- Ask them to explain their reasoning.
- Compare alternatives with a mentor.
- Revisit the outcome; expand responsibility.
More people can exercise judgment next time.
From there, the learning belongs in real tasks. Possessing an instruction and knowing how to repair it are different abilities.Brandon Gell describes a marketing client that wanted prompts for email copy. Midway through, Every rebuilt the engagement around working sessions (opens in a new tab): how the prompts were designed, why they worked, and how to adapt them. That detail makes the capability argument tangible. The deliverable could work today and still leave the team dependent tomorrow. This is a practitioner’s account of one engagement, not a controlled comparison of training methods.
I like working alongside someone while they try a workflow, explain what they notice, and change one thing. Their questions improve the workflow and the teaching. Choosing a model, giving useful context, and coordinating parallel efforts all become more concrete when there’s actual work on the table.
Champions can make today’s good work visible through show-and-tell, lunch-and-learn, mentoring, and kudos. A time-bound hackathon can put a real problem on the table and make experimenting fun. Catering doesn’t solve overload; leadership still decides how learning and teaching fit the workload. Promotion, compensation, and bigger opportunities can reward maintenance and teaching alongside the impressive first demo.
For tomorrow’s experts, I preserve decision practice: ask a learner to propose a course of action, explain the evidence, compare alternatives with an experienced colleague, and revisit what happened. Give them increasing responsibility with appropriate supervision. A prompt library can support that work; it cannot supply the experience of making and revising a judgment. Which decisions best replace the lost apprenticeship tasks is a question I’m actively exploring.
Adoption and support
If this is you: the tool works, the champion uses it every day, and everyone else quietly went back to the old way.
A useful tool with low use calls for diagnosis before persuasion. I watch the attempted workflow and ask where the person returned to the old way. Is the missing piece coverage, habit, discoverability, installation, support, trust, or a result that creates more checking? Non-use is evidence to investigate, not a personality flaw.
Changing how people work can be harder than building the technology. I plan for that from the start: time to try the workflow, help when it breaks, a way to report friction, and someone responsible for the adoption effort.At the September 9, 2026 Developer Marketing Summit (opens in a new tab), Gal Bakal gave a talk on building and rolling out an internal developer platform at Yahoo. He described two days of building and six months to reach 700 users. The intervening work included simpler installation, a community channel, updates inside the tool, and responsive support. Later, the package became part of onboarding for developers entering the AI transformation program. Even among engineers, a useful tool needed a sustained adoption effort. That is the planning lesson I take from his account: change management needs an owner and time of its own. The adoption milestones are his reported experience, not an independently measured business result.
Choose the repair by the signal: repeated installation failures call for support; repeated corrections call for better guidance; missing task coverage may mean the tool is a poor fit. When the workflow works but colleagues keep forgetting it, I’d put a supported first attempt into an existing work moment. Then check repeat useful completion, not just the first try.
My next move would be to pair a champion with a support owner on a recurring frustration: make the gap easy to report, review a repair, and check that the person can use it. Usage data helps diagnose fit and support. I still want evidence of useful completion, quality, and total effort before calling it an outcome.Gal Bakal’s Knowledge Activation paper (v2), sections 8.3–8.8 (opens in a new tab), describes his Yahoo deployment and a voluntary, single-arm survey of largely AI-fluent respondents. Benefits are self-reported, with concurrent changes limiting causal attribution. I find the reported friction particularly instructive: coverage, discoverability, habit, and stale guidance each suggest a different repair. Positive sentiment can help identify something worth investigating; it does not by itself measure business impact.
Follow the result into someone’s day
If this is you: the drafts arrive faster and the decision still waits three days, or the brief is beautiful and the reviewer’s hour got longer.
A generated artifact can be accurate and still make the job harder. I ask which existing step it replaces, how it should be read, and whether it adds another review obligation. Preparation time saved can disappear into checking and reconciliation.
I evaluate total effort across the workflow, including review, rework, and maintenance. I also watch for the next queue: more drafts can leave the same decision-maker with a larger pile. Useful completion is the measure that matters here.
Where did the saved time go?
What happens to preparation, review, rework, and maintenance together?
The effort moves elsewhere
More checking. More waiting. A larger pile for the same decision-maker.
Next move: repair the constraint that absorbed the gain.
Useful capacity is released
Leadership gives that capacity a destination.
- Deeper customer work
- New experiments
- Quality or breathing room
- A specific cost decision
I keep observation open after initial testing. A passing evaluation and a person’s working day provide different kinds of evidence.In his conversation with Dan Shipper (opens in a new tab), Notion cofounder Simon Last describes finding whole new classes of errors after building substantial evaluation sets. Around 28:53–32:14, he explains the value of collecting, labeling, and retesting failures from use. What stays with me is the discovery problem: the evaluation can only cover errors someone has thought to look for. His account offers no universal reliability target.
I make it easy to report a miss, find who owns it, and see whether it was resolved. I also retain contact with customers and the physical work of the business. A synthesis of feedback can help direct attention; it doesn’t substitute for hearing what happened from the people involved.
I establish a baseline before the workflow changes and choose measures that inform a decision. Early signals might include completed work at an agreed quality, fewer repeated corrections, or shorter time to a useful customer response. Later measures might include conversion, operating cost, or customer uptake of a new offer.
I want business outcomes on the leadership scorecard while keeping attribution difficulty visible.Lauren Morgenstein Schiavone advocates business-outcome measurement in this discussion of AI transformation (opens in a new tab), while acknowledging attribution challenges. I use that practitioner distinction to keep the scorecard ambitious and the causal claims proportionate to the evidence. An improved result may coincide with changes in demand, staffing, or pricing. A before-and-after comparison can be useful without proving AI caused the difference.
Customer support makes the distinction concrete. Deflection means an interaction avoids or moves away from a staffed channel; resolution means the customer’s problem is actually solved. A customer who gives up can disappear from the queue without getting help. I inspect repeat contacts, quality, escalation, and rework alongside speed.The Enterprise AI Playbook’s measurement appendix, page 108 (opens in a new tab) separates resolved interactions from apparent deflection. That is a useful definition for evaluating a support workflow. It does not establish an effect size or tell us that any particular reduction in staffed contacts means customers received better help.
I separate five observations before calling something a financial return:
- Released time: less effort for useful completion at an agreed quality. This is capacity, not cash.
- Redeployment: that capacity goes to named work, with the skills and demand to use it.
- Avoided costs: a planned hire or expense is no longer needed. The counterfactual needs explaining.
- Actual costs: spending really changes, after model use, integration, review, transition, and ongoing maintenance are included.
- Revenue: customers buy or retain something valuable. More activity and a larger pipeline are intermediate signals.
Then I reconnect the measurement to the people using the work. Customer conversations and operational observation can expose something a dashboard misses. If faster responses become less useful, the speed improvement needs reconsideration. I want the feedback to change the workflow, including the decision to stop it.
Make ownership something a person can demonstrate
If this is you: the consultant left a folder, the workflow broke the first time a model changed, and nobody inside can fix it.
I involve the internal owner while choices are still being made. They need the reasoning behind the workflow and the authority to keep it useful. Receiving a folder at the end leaves too much of that understanding outside the organization.
My handoff standard is practical: the owner makes and tests a real change, then completes a recovery exercise. Can they notice a problem, pause the work where needed, and restore a known working version? Documentation helps them do this; the demonstration shows where more support is required.
I finish by making delivered work, deferred work, and accepted responsibilities explicit. I review observed outcomes with their limitations and identify the next constraint. Leadership decides how much capacity goes into continued improvement and how people share in the upside. An experiment has an end date. A useful capability needs someone to care for it afterward.
A note on the evidence
The research and practitioner accounts in this guide help me ask better questions. I use them to sharpen situational judgment, and I keep the limits beside the claims they qualify.Elisa Pereira, Alvin Wang Graylin, and Erik Brynjolfsson’s Enterprise AI Playbook (opens in a new tab) studies 51 selected successful projects in 41 organizations. That selection is useful for examining what happened inside promising deployments. It cannot tell us how often an approach succeeds across all attempts. The accounts are qualitative, mostly from executives and project leaders, with varied outcomes; they are neither independent employee research nor proof of what caused success.
I’d begin with one piece of work someone already cares about. Follow it all the way through, including the person who has to use the result. There’s usually a better starting point there than in a catalog of AI features.
Read: The work after the work gets faster · Compare notes with me