Somewhere in a system I built there is a description of a man named Dave. It says he escalates under stress and needs the context up front or he digs in.
It’s accurate. I wrote it, more or less, by asking for help dealing with him.
Dave has no idea.
I’ll come back to Dave. First, what I was trying to build when I wrote that down about him.
An outcome worth having
Picture a thirty message email thread. Six people, four days, two of them arguing past each other, one attachment that matters, and a decision buried somewhere around message nineteen.
Today you have two options, and neither one gets you out of the house.
Option one is to read the whole thing. Not because reading it is valuable, but because you can’t know which four sentences matter until you’ve read the other three hundred. Ninety minutes reconstructing context so you can make a thirty second decision.
Option two is to have an AI summarize it. And it will do that well, in the sense that the summary will be short and accurate and would satisfy any reasonable reader.
That’s exactly the problem. It picked out what matters to a reasonable reader. It doesn’t know that Dave’s throwaway line in message twelve is the tell, because Dave only writes like that when he’s about to escalate. It doesn’t know the decision in message nineteen collides with something you committed to in a different thread last month. It doesn’t know which two of these six people you’d drop everything for.
So you skim the thread anyway, because you can’t see what it left out. Option two cost you a step instead of saving you one.
A summary you have to verify is worse than no summary at all. And you have to verify anything that doesn’t know you.
Now suppose there were a third option. Something reads it and comes back with: here’s the decision they need from you, here’s the one line that makes it make sense, Dave is about to blow up and you should call him, the rest is noise and I filed it.
You did not save eighty-eight minutes to sit and do nothing with them. You went to your kid’s soccer game.
And you called Dave from the car on the way.
That is the whole thing. Not less time working. The same time, spent on what you would actually choose. The value was never the minutes, it is the substitution of what the minutes went to.
The usual objection here is that saved time always gets absorbed, and the historical evidence for it is strong. Ruth Schwartz Cowan’s More Work for Mother documented that decades of labor saving appliances left total household labor hours roughly flat, because standards rose to fill the space. True. Also beside the point. Those machines still eliminated hauling water and boiling laundry over a fire.
Flat hours. Transformed hours. I’ll take that trade every time.
Why nobody builds option three
Here is where I think most AI assistants quietly fail, and the failure is disguised as success.
Everybody ships option two. Option two is safe, and it’s safe precisely because it never has to be wrong. It shortens the pile and hands the judgment straight back to you, and the industry calls that respecting the human in the loop. What it actually respects is the vendor’s liability.
Option three is a different act. It means taking a position. This needs you. This does not. This is handled. Call Dave.
Nobody wants to build that, because the moment you take a position you own the misses.
But there is no other route to the soccer game. You don’t leave the house because the summary was well written. You leave because you believe nothing is on fire. That belief is the actual product. Everything else is scaffolding underneath it.
And a belief you can’t check is just advertising.
So it has to show its work. Everything it did for you, and more to the point everything it decided not to bring you, written in language you’d actually use instead of a log file. Here’s the part I think most people get backwards though: you should never be obliged to look. A system that’s only safe when you audit it isn’t safe. It has just moved the work to your side of the desk and called it transparency.
The record is there so looking is possible and cheap. Not so it becomes one more thing on your list.
There’s a selfish reason for it too. The pile of things it decided you didn’t need to see is precisely where its mistakes are hiding. Somebody glancing at that pile and saying “no, that one mattered” is the best correction the thing will ever get.
What we aimed it at
I’ve written before that your AI strategy is a labor strategy, and I stand by it. Short version: we aimed this at headcount because headcount was the thing with a line item, not because it was the best available target. The numbers are in the addendum, including the ones that complicate it.
One number I won’t leave back there. Stanford’s AI Index has 73% of AI experts optimistic about AI, against 23% of the public. Fifty points.
When the people building a thing are that much happier about it than the people receiving it, that isn’t a communications problem.
What I want to talk about is what else it could be aimed at.
Personal, not personalized
Something changed underneath all of this. Back in February I wrote about what happens when software gets cheap enough to build for one person. That constraint is gone now. Software for a single person costs an afternoon.
Which means the era of software built for workforce productivity is ending, and an era of software built for individual productivity is starting.
And the word that matters is personal, not personalized.
Personalized software is a product built for a segment with your preferences painted on top. Your name in the header, your timezone, dark mode. Personal software is built for you, adapts to you, and answers to you. Not to a persona, a demographic, or a role.
I want to be careful here, because when people hear personal they hear small, and that isn’t it at all. This is a distinction about who the software is designed for, not about how many people use it. A product can serve a million people and still be personal, if every one of that million gets something genuinely shaped to them. A product can serve exactly one person and still be merely personalized, if it’s a generic thing wearing their preferences.
The same distinction just renders differently at each end. For an audience of one it renders as construction, because building the thing around the person now costs an afternoon. For an audience of a million ones it renders as calibration, where the software is shared but the memory, the judgment about when to interrupt you, and the sense of whether it actually landed are yours alone.
What can’t survive either way is the average user. The moment you design for the median of a population you’re back to personalized, however many preference toggles you ship.
The average user doesn’t have a Dave. You do.
And belief is earned one person at a time
The industry has settled on observability. Watch the model, score the outputs against a rubric, aggregate across all users, chase the benchmark. Useful discipline. It answers a third person question: is the model performing well?
That is not my question. Mine is second person, it has an N of one, and it has no ground truth anywhere outside one specific human being.
Did I get this right for you?
Do I know Dave gets an answer inside the hour and the vendor rep can wait until Thursday? Do I know this is the thread you’ve been dreading, and that “resolved” matters to you more than “answered well”? A benchmark score is a personalized answer. Correct for the average of everybody, which is to say correct for nobody in particular.
This is why memory, learning, and adversarial quality review are not features I bolted onto the side of my system. They are the system. Without memory, “right for you” has no referent at all. And the evaluation can’t run somewhere far away against a generic rubric, because judging whether something landed for you requires knowing you. I used to justify keeping that local as a privacy property. It isn’t. It’s a correctness property. Privacy is just what you get for free when you build it correctly.
What I’m going to do about it
I’m open sourcing the harness underneath KAIPARR.
Not the app. The thing beneath it: the runtime, the secure execution on your own machines, the memory, the calibration loop, and the governed improvement cycle where nothing changes without a human approving it, and anything that goes wrong rolls back.
Here’s why it has to be open. Writing personal software now costs an afternoon, but the parts that make it trustworthy do not. Memory that persists and stays yours. Calibration that works out Dave is worth interrupting you for and the newsletter isn’t. Safe execution against your real accounts and your real files. Connections to where your life actually happens. Nobody rebuilds that from scratch, whether they’re serving one person or a million.
So either that infrastructure exists as shared, inspectable plumbing, or the era of personal applications turns into a handful of vendors renting personhood back to us by the seat.
It’s a harness and not an app for a reason. An application encodes one answer. A harness is an instrument for finding one. And if somebody else uses it to find a better answer than mine, that isn’t a loss, it’s the entire point.
You’ll bring your own model credentials, so you own the relationship with Anthropic or OpenAI or whoever, not me. Your data will not be the product and will never be training material. If you want to leave, you leave with all of it.
The part I can’t tell you
Here’s what I don’t know, and I’d rather say it than pretend.
I don’t know how you build software that genuinely adapts to each person while still serving a shared outcome. Not really. I know what it isn’t. It isn’t configurable forms and admin screens, where a vendor guesses in advance at every way you might differ and hands you a menu. And it isn’t stuffing a profile blob into a prompt and hoping, which is most of what ships today and which nobody can version, measure, or explain.
I have a hypothesis. The application declares the outcome and the agent works out the path. What varies per person, meaning what matters to you, when you want to be interrupted, how it should sound, what it’s allowed to do without asking, becomes a real versioned artifact the system writes from watching you, rather than a settings page you fill out. And it gets governed like any other change, with approval and rollback, because a thing that adapts to you without you being able to see it or stop it is a different and worse product.
I might be wrong about all of that.
There are problems underneath it I can’t currently solve. You can’t A/B test one human being, so I don’t know how to prove a change helped you specifically. A new user has no history, so for the first few weeks it can’t tell Dave from the vendor rep, and personal has to be earned. Learning from a population makes each individual converge faster, which is useful and also means using other people’s behavior, and I don’t yet know exactly where I’d draw that line.
None of that gets solved in a library or a benchmark, which is why KAIPARR the product isn’t going anywhere. It stays in production, with real users, real data, real scheduled jobs, and real runners on real machines. How long does it take to actually know someone? Can you tell whether a change helped one specific person without an experiment you’re not able to run? Those are questions about human beings over months, and they need people doing real work with real consequences when the thing gets it wrong.
So the product is the lab. That’s not dogfooding, it’s the only instrument that exists for this.
Which puts an obligation on me, and I’d rather write it down where you can hold me to it. If you use KAIPARR, you’re a participant and not a subject. You should know it’s learning. You should be able to look at what it thinks about you and tell it when it’s wrong. And whatever it figures out should show up in your Tuesday afternoon before it shows up in anybody’s pitch deck.
The thing I’m actually afraid of
It isn’t that this doesn’t work.
It’s that it works, and becomes another feed.
Sit with what I’m describing for a second. Something that holds years of your context. That knows what you care about, who you answer inside an hour, when you’re reachable, and what you’ll actually act on. Everything that makes that good at serving you makes it the most precise attention extraction engine anybody has ever built. The loop that learns when to interrupt you for your benefit is the identical loop that learns when to interrupt you for somebody else’s.
We have watched this happen before, in slow motion, and most of us are still carrying it around in our pockets.
Facebook and Instagram and everything after them started out doing something people actually wanted. Show me what my friends are up to. Show me more of the stuff I’m into. Nobody in that first meeting set out to build a slot machine. But the thing that made the feed good at knowing you is the identical thing that made it good at holding you, and once the revenue depended on holding you, “what you’re interested in” quietly turned into “what keeps you here.” The technology was never the problem. The business model decided what the technology was for.
Now picture that same arrangement with something that never had to guess.
A feed infers you from the outside. It watches what you clicked, what you lingered on, what you scrolled past, and assembles a model of you out of behavior you never meant to hand over. What I’m describing doesn’t infer much of anything. You told it. Every question you asked and the way you phrased it. Every draft you rewrote before sending, which is you saying precisely, this is not how I sound. The times you got short with it, which is the clearest signal that exists about what you actually care about. The one thing it did that you liked enough to go mention to somebody. Years of that, sitting on top of all the mail and messages and calendars it was reading in order to help you in the first place.
The feed had to work for its model of you. This one gets handed the answers, by you, in the moments you had the least reason to be guarded.
And it’s worse than that, because it was never only about you.
You told it about Dave. That he escalates when he’s stressed, that he needs the context up front or he digs in. You told it which colleagues you trust and which ones you route around. You told it your sister is going through something and to flag anything from her. Not one of those people agreed to any of this. They were talking to you. A system was listening, and it now holds a model of them built from the most candid description of them that exists anywhere, which is the one you gave in private while asking for help.
So if a second party ever gets into that room, you haven’t just sold yourself out. You’ve sold out everybody who ever trusted you enough to be a person around you.
And people can smell it coming. Ipsos asked across 32 countries whether they’d trust a generative AI tool less if its answers were influenced by advertisers. 46% said yes, and in 15 of those countries it was an outright majority. Nobody needs this explained to them. They’ve lived it.
So the difference isn’t technical, it’s about who else is in the room. There’s one party this thing answers to, and it’s the person using it. No ads, no placement, no sponsored anything, no affiliate nonsense, no selling aggregate behavior to somebody who wants to reach you. And I’m not holding that door open for later, for when the numbers get hard and somebody sensible explains that one small tasteful ad unit wouldn’t really hurt anybody.
But I want to be straight with you about something, because a promise from me is worth roughly as much as I’m around to keep it.
I’ve watched enough companies get bought. The founder means every word, right up until the founder isn’t the one deciding anymore. So I went looking for the version of this where it isn’t up to me at all. Encrypt everything with a key only you hold. Then whoever owns this in ten years can’t sell your interior, not because they promised not to, but because they mathematically cannot. No terms of service update gets around arithmetic.
I hit a wall.
If only you can read it, then every part of this that has to think about you has to run on your phone. The evaluation. The learning. The work that happens overnight. And your phone sleeps, throttles anything running in the background, and can’t run the model at all. The whole point of the thing is that it works while you aren’t looking.
A phone doesn’t work while you aren’t looking.
So I don’t have it. Not yet, possibly not ever in the clean form I wanted, and I’d rather say that flatly than let you assume I’ve solved something I haven’t. This is the single hardest problem in what I’m building and I’m not going to pretend it’s a detail.
What you get instead is weaker, and here it is honestly. My word, which I mean, and which is worth exactly what everybody’s word is worth. Code you can read, because it’s open source, so any betrayal would at least have to happen in daylight. Constraints written down in public where you can quote them back at me. And one number.
I’ll hand you that number so you can hold me to it. Watch the interrupt count. I’m building this so that number goes down over time. Every attention business on earth needs it to go up. If you ever catch me celebrating engagement, daily actives, or time spent, I’ve lost the plot and you should stop trusting the thing.
So why build it at all.
Because it’s getting built regardless, and not because of me. The version that harvests you is the easy one, fifteen years of prior art and a business model that already works, and nobody needs my help with that. The one actually pointed at you is what doesn’t get built by default. I’d rather it exist and get argued over in the open than not exist, and I’d rather the person building it be asking these questions now, while the answers can still change the code, than after fifty million people are already inside it.
That isn’t a guarantee. It’s a starting position, and it’s the only honest one I have.
What I want is small. Somebody closes a thirty message thread in ninety seconds instead of ninety minutes, calls Dave because it turned out that mattered, and gets to the field before the second half. None of that works unless something knows Dave.
You can consent to being targeted. You cannot consent on Dave’s behalf.
Nobody asked Dave.
Addendum: I didn’t make this up, and here’s what argues with me
You don’t need any of this to get the point. It’s here because I make claims above that somebody should be able to check, and because the evidence that complicates my argument deserves to be visible rather than buried.
On aiming it at headcount. Goldman measured roughly 16,000 net US jobs lost per month to AI displacement, concentrated in clerical and high-substitution roles. Stanford’s AI Index found employment for software developers aged 22 to 25 down nearly 20% since 2024, while developers over 30 kept gaining. So the effect lands hardest on people trying to get in.
On the returns not showing up. A 2026 NBER survey found 89% of executives reporting no measurable productivity impact, despite roughly 70% of firms using AI. Task-level speedups are real. They keep getting absorbed by coordination overhead nobody automated.
On trust. Gallup has 27% of Americans trusting businesses to use AI responsibly, down from 31%, with distrust rising for the first time since they started asking. Over the same period roughly 52% of US employees say they use AI at work, about double two years ago. Half the country is using a technology it doesn’t trust.
On advertising specifically. Ipsos asked across 32 countries whether people would trust a generative AI tool less if advertisers influenced its answers. 46% said yes, a majority in 15 markets.
What cuts against me
The global mood isn’t hostile, it’s ambivalent. Across 25 countries the median is 34% more concerned than excited, 16% more excited than concerned, and 42% feeling both at once. On several measures optimism has been climbing, not falling.
The pessimism is regional. Roughly three quarters of people in the Middle East and Africa see AI as a tool for growth and empowerment, against under a third in North America and Western Europe. 83% of Chinese respondents are optimistic. 39% of Americans are. The gloom I’m describing concentrates in the wealthy Anglosphere, which happens to be where most of this technology is built and sold. Read that either way you like.
People who use it heavily report real gains. Among those using AI across many tasks, a large majority report a positive productivity effect. Disillusionment and genuine utility are coexisting, which is harder to write about than either one alone.
And I’m not a neutral party. I’m building one of these things. Weigh accordingly.