Make AI Adapt to Human Diversity — Not Humans to AI
AI is learning what good looks like. It hasn't learned what good looks like for you.
Ask two people to describe a great assistant and you will get two incompatible answers.
One wants it to just handle things. Book the flight, file the expense, send the follow-up — don't ask, don't explain. When it works, they say great, it just handled it.
The other wants to be consulted. Show me what you found before you decide. Confirm before you spend anything. When that person meets the first person's assistant, they say wait — it paid without asking me.
Same agent. Same action. One person delighted, one person's trust permanently damaged.
People differ in risk tolerance, in how much control they want, in speed versus thoroughness, in privacy expectations — even in what counts as a satisfactory outcome. Today's AI compresses all of that into one universal optimal behavior, trained on averaged preferences. Everyone gets the same agent, and humans bend themselves around it.
The axis that diverges
There is a distinction hiding here. Almost all preference data collected today measures what an agent produces: which output is better. That is outcome taste, and it converges toward one shared answer. Ask enough people which of two summaries is cleaner and they will largely agree.
Procedural preference is a different axis entirely. It is how the work should be done: when to ask versus act, how much to verify, what to do when something is ambiguous, whether to show reasoning or just deliver. Unlike outcome taste, it diverges by person. And it is exactly what agents acting on your behalf need to know.
The idea isn't new. Economists named it in 2004 — procedural utility, the finding that people care how an outcome was produced and not only what it is. What's new is that agents finally make it observable.
Two axes, not one ladder
At first it looked like a sequence — correctness, then taste, then personalization. It's a tidy story and it's wrong. It collapses two independent questions. One is what gets evaluated: the output, or the way the work was done. The other is whose preference counts: everyone's, or yours.
Cross them and you get four combinations, not three stages.
Three of the four are being worked on. Averaged outcome preference is collected by head-to-head comparison — people vote on which design or answer is better, and the votes aggregate into a ranking. Individual outcome preference is what recommendation systems have worked on for twenty years — your aesthetic, your topics, your tone. Averaged procedural preference is what alignment research asks about: when should any agent pause before acting?
The fourth is individual procedural preference — how you want work carried out. It is the one we know least about, for a structural reason: it only became observable when agents began taking real actions on our behalf, which is a development of the last two years. You cannot study taste in action until there are actions to watch.
Personalization research has grown quickly in that time, but nearly all of it aims at what a model produces — the response, the tone, the recommendation — rather than how a task gets carried out.
Difference is the data
The instinct in machine learning is to treat variance as noise to be smoothed. Here it is the opposite. The spread between people is the signal — it is the whole thing being measured. Average it away and you get an agent that is acceptable to everyone and right for no one. Todd Rose made this case about cockpits and classrooms — design for the average person and you've designed for nobody. Agents are where it stops being an argument about ergonomics.
The signals already exist, and they are behavioral rather than declared. A person edits an output, overrides an action, takes control midway, redoes the work, or abandons the agent entirely. What people actually do says more about how they want to be served than anything they would type into a settings panel.
Today those moments are flattened into clicks or buried in execution logs. The builder sees only that something failed — not whether the user wanted more control, a different sequence, greater thoroughness, or simply a confirmation before the agent acted.
It is a kind of segmentation, but not the kind anyone knows how to do. The old version was demographic and static: age, industry, company size. What's observable now is procedural — not who you are, but how you work. Not enterprise customer in financial services but verifies before committing, wants sources shown, never delegates payment authority.
An end to averaging
Human diversity should be represented inside AI systems, not averaged away.
Agents built this way take over the parts of work people genuinely want to delegate, while preserving control, individuality, and human agency. The long-term shape of it is a portable memory of how you want work done — collected from your real behavior, owned by you, and carried into any new agent, product, or robot so it immediately knows how to work with you.
That ownership matters. If personal context ends up locked inside proprietary platforms, each assistant hoarding its own private model of you, we get personalization without agency — and switching costs measured in years of accumulated understanding.
That is the bet. Not a smarter average, but an end to averaging.
Idios Labs is building the behavioral-intelligence and personalization layer for AI agents. If you run an agent product with a consequential workflow — or you think about this problem too — I would like to hear from you.
hello@idioslabs.com