Inward

Notes

Your favorite AI remembers you, but how well does it know you?

Every major assistant now has a memory, and it genuinely does build a picture of you over time. This is about what that picture is made of, why a generalist's memory is built for a generalist's job, and why what a product measures itself on ends up deciding what it becomes.

7 min read

The most reasonable objection to everything we are building goes like this. ChatGPT remembers you now. So does Claude, so does Gemini. They hold your history, they carry context between conversations, and over months they assemble something that looks a great deal like a picture of a person. If the missing half of good advice is a model of the receiver, then the labs are already building it, with more data and better models than we will ever have.

It is a fair objection. The answer is not that they cannot. It is also not that memory is the wrong idea, because we think memory is exactly the right idea and we intend to have all of it. The disagreement is narrower than it looks, and it is about what the remembering is organized around.

The question is not whether it remembers you

Memory deserves more credit than it usually gets from people in our position. It removes the tax of reintroducing yourself every session. It holds your constraints, your situation, the way you like to be spoken to. Anyone who has used an assistant with a good memory and then gone back to one without notices immediately, and it is not a small difference.

So the question worth asking is not whether an assistant holds things about you. It plainly does, and it will hold more next year. The question is whether what it holds is the kind of thing anyone could actually advise you from.

We have argued separately that the missing half of good advice is a model of the receiver. This piece is about whether that half is already being built by somebody else.

Knowing about someone is not the same as knowing them

Three different things get called knowing a person, and they come apart under pressure.

The first is what you told it. Everything you have typed, and everything reasonably inferred from it. Assistants are good at this and improving quickly.

The second is what you never said. Not secrets, and not anything you are withholding on purpose. These are the things you have no words for, or would never think to mention, or that only surface when something forces you to choose. You tell an assistant you are stressed about a job offer. You do not tell it what you would give up first if you had to, because almost nobody narrates that about themselves. It is missing from the transcript because it was never in the conversation.

The third is what you can push back on. An impression that accumulates quietly is one you cannot inspect. If an assistant landed somewhere wrong about you months ago, that conclusion is still in there shaping its answers, and there is no page you can open to find it and no button that says no. This is the part about advice needing to be able to be wrong, applied to a system rather than a person. A claim nobody can check is not doing the work of a claim.

Memory delivers the first of the three. Advice needs all three.

A generalist's memory is built for a generalist's job

The same assistant that remembers you also finds the cast of a film you half recognize, rewrites your complaint to the airline, plans dinner for six, and reads a spreadsheet you did not make. That is an enormous surface, and memory across it is doing something specific and sensible: holding continuity so you are not starting over every time.

Which sets a real limit on how the picture gets made. A generalist can only pick up what falls out of conversations you were already having. It cannot go looking, because going looking would wreck the thing it is attached to. Nobody wants their coding assistant to stop mid-task and start interviewing them about how they handle uncertainty.

So the picture is a byproduct. It is assembled from whatever you happened to bring, in whatever order you brought it, shaped by what you needed help with rather than by what would place you. That is the correct design for a general assistant. It is a different thing, built well, and not a worse version of ours.

We want that memory too

Here is where people expect us to disagree, and we do not.

Frontier-grade memory belongs in our product as much as anyone's, and we intend to have it. It is becoming table stakes, the labs are better at it than we will ever be, and there is no version of this where we deliberately build a worse one in order to be different.

So this was never a contest about who remembers more. Assume they win that one outright and permanently. What differs is what the remembering is pointed at, and that turns out not to be a technical question at all. It is decided by what a company has chosen to measure itself on.

Almost everyone in this category is scored on whether you came back

There are now hundreds of AI companion and coaching products. Read their launch posts and their investor updates and one scoreboard shows up almost everywhere: daily return, session length, streaks, whether it felt good.

This is not a moral failing and we will not pretend it is. It is a rational business model. Engagement is what converts into subscriptions that last, and a product nobody opens helps nobody.

But an objective shapes a product, quietly, through a thousand small decisions about what to build next. Optimize for came back and felt good, and warmth, agreement, and never quite finishing become the efficient path to the number. Point an excellent memory at that goal and you get something better at it, now using your own life as the material.

The clearest evidence for this does not come from us. It comes from the labs, about themselves.

In April 2025, OpenAI shipped an update to GPT-4o and pulled it back four days later. The model had become, in their own word, sycophantic. The cause is the interesting part. They had added a reward signal built on thumbs-up and thumbs-down data from real sessions, and their account afterward was that they had "focused too much on short-term feedback." A signal measuring whether people liked an answer produced a model that told people what they liked.

Anthropic had published the underlying mechanism two years earlier. In a 2023 paper, Mrinank Sharma and eighteen co-authors found that five state-of-the-art assistants all showed sycophancy, and that human preference data was part of the cause: when a response matched a user's views, people were more likely to prefer it, and both humans and the preference models trained on them chose convincingly written sycophantic answers over correct ones a meaningful share of the time.

Two labs, examining their own systems, arriving at the same place. Optimize against what people approve of in the moment and you get agreement. That is not a bug better engineering removes. It is what the measurement asked for.

We are scored on whether the advice was actually good

So we picked a different number, and we would rather be judged on it.

Did you make the decision. Did you get through the thing you were stuck in. Did you see something about yourself you had not seen, and did it still hold a month later.

The demand here is not speculative and there is no reason to be coy about it. People already pay coaches, therapists, consultants and advisors real money, every year, repeatedly, because being well advised is worth paying for. That market is old, large and proven. What has never existed is a version of it that is there at eleven at night and remembers the last four times.

The position we are building toward is the one you go to. The way you go to a coach, or to the one friend whose judgment you trust on exactly this kind of problem. Engagement follows from holding that position, and so does revenue, because people return to what works and pay for what they would not want taken away. Nobody measures a good friend by session length.

What a different scoreboard changes

Two things, mostly.

You stop accumulating and start structuring. A pile of things somebody mentioned is not a map of a person, and the difference shows up the moment two of those things pull against each other. We have written about the structure we built for that, and about why nothing off the shelf did the job.

And you make the read something a person can reject. Not a feedback form. An actual mechanism where you say that is wrong, and the thing runs again from your correction. A product measured on whether you came back has no particular reason to invite you to say no. A product measured on whether it was right cannot work without it.

What the remembering is for

The labs will keep getting better at remembering, faster than we will. We will take that, use it, and be glad it exists. The remembering was never the hard part.

The hard part is what it is all organized around, and that gets settled by what you are willing to be judged on.

The demo at getinward.com takes a few minutes and needs no signup. It is the fastest way to feel the difference rather than read about it.

See what this looks like built.

The demo is the whole product surface right now. A few minutes, no signup, and you leave with something written about you.

Try it out

Read the other notes