This product discovery case study started with what looked like a simple ChatGPT search problem: users knew they had created something valuable before, but returning to it was often harder than expected. What began as a search question eventually turned into a much broader investigation into how people rediscover, recover and resume AI-assisted work.
It started with a very ordinary frustration.
You know you discussed something with ChatGPT before.
Maybe it was a product idea. A SQL query. A piece of research. A decision you made after going back and forth for twenty minutes.
You remember the work.
You just don’t remember where it is.
So you search.
You try another keyword.
You open three chats that look vaguely familiar.
Wrong one.
Another search.
Another chat.
At some point, a strange thing happens.
You think:
Forget it. I’ll just ask again.
That was the starting point for this product discovery exercise.
My first instinct was obvious: this looks like a search problem.
Maybe ChatGPT needs better search.
But that answer came far too quickly.
And when an answer comes that quickly in product work, it is usually worth asking whether we have understood the problem at all.
So instead of designing a search feature, I started investigating.
First question: is this even a real problem?
A few complaints on Reddit are not enough to justify a product decision.
So I started with social listening.
I looked through discussions from people who use ChatGPT heavily: researchers, developers, writers, people managing projects, people with hundreds of conversations in their history.
The complaints were remarkably familiar.
People remembered having useful information somewhere, but not the exact words they had used.
Some had so many conversations that browsing them had become difficult.
Others could find the correct conversation but still couldn’t find the particular answer buried inside it.
And some simply started a new conversation because recovering the old work took too much effort.
That was the first important moment in the discovery.
The problem wasn’t behaving like one problem.
It was behaving like several.
The people feeling it most weren’t necessarily “all ChatGPT users”
Someone who asks, “What’s the capital of France?” probably doesn’t care whether that conversation is easy to retrieve six months later.
But imagine someone using ChatGPT for a product strategy over three weeks.
There might be separate discussions about:
customer research,
pricing,
positioning,
competitors,
MVP scope,
metrics,
and a dozen decisions made along the way.
That work has future value.
And that led me to focus on a more specific type of user:
people who use ChatGPT repeatedly for knowledge work and expect to come back to previous work later.
The number of conversations matters.
But something else matters even more:
whether those conversations contain work worth returning to.
That distinction turned out to be much more useful than simply calling them “heavy users.”
Then I mapped what actually happens when someone comes back
Imagine you are trying to continue something you worked on two weeks ago.
The journey looks roughly like this:
You remember that the work exists.
Then you try to find the conversation.
Then you have to work out which of several similar conversations is the correct one.
Then you open it.
Now you have another problem.
Where, inside 80 messages, was the thing you actually needed?
You scroll.
You read.
You eventually find the answer.
But you still aren’t necessarily ready to continue.
What had already been decided?
Why did you reject the other option?
What constraint did you discover?
What were you supposed to do next?
Suddenly, finding the chat looks like only one small part of the problem.
The actual journey was closer to:
Remember → Find → Recognise → Recover → Rebuild context → Continue
And this was where my original “better search” idea started falling apart.
Search wasn’t the whole story
Once the journey was visible, I started forming hypotheses.
Maybe people cannot remember exact keywords.
Maybe they remember the meaning of the conversation rather than the language used inside it.
Maybe similar titles make chats hard to recognise.
Maybe long conversations make useful answers expensive to recover.
Maybe the user finds the answer but still has to reconstruct too much context.
And maybe there is a simple behavioural threshold underneath everything:
When recovering previous work feels harder than recreating it, people start again.
That hypothesis interested me the most.
Because if it is true, this isn’t just a navigation inconvenience.
It means valuable work can effectively become disposable.
The discovery changed the question
I began with this:
How could ChatGPT make old conversations easier to find?
By the middle of the exercise, that question no longer felt right.
A better question was:
How could ChatGPT make valuable previous work easier to continue?
That sounds like a small wording change.
It isn’t.
“Find an old chat” immediately pushes the product team toward search.
“Continue previous work” opens a much bigger opportunity space.
Now you can think about three separate jobs.
Rediscover
Which previous work am I looking for?
Recover
Where is the useful part inside it?
Resume
What happened before, and where do I continue?
That became the real structure of the problem.
Only then did I allow myself to think about solutions
Once the opportunity was clearer, I explored several directions.
One was natural-language retrieval.
Instead of remembering the exact title, a user could say:
“Find the conversation where I was comparing product management and analytics careers.”
Another direction was better navigation inside long conversations.
Another was automatically surfacing important decisions and outputs.
Another was something like a “Where you left off” state when reopening old work.
And another was treating previous conversations less like a list of chats and more like a personal body of work that could be queried.
The interesting part wasn’t choosing the biggest idea.
It was figuring out what assumption needed to be tested first.
Because even a very impressive solution is still just a guess until somebody uses it.
The concept I would test
The concept I eventually narrowed down to was something I called Work Recall.
The idea is simple.
Instead of asking the user to remember the conversation, ask them to describe the work.
For example:
“Find the work where I was researching why people struggle to retrieve old ChatGPT conversations.”
The experience could return a few possible matches, but with more context than a title alone.
Not just:
Product Discovery
but something closer to:
ChatGPT Retrieval Discovery
You were investigating why frequent users struggle to return to old AI-assisted work.
Topics discussed: social listening, long conversations, retrieval, context reconstruction.
Last meaningful decision: the problem appears broader than search.
Then, after opening it, the system could help answer another question:
Where did I leave off?
That is the part I find most interesting.
Because the product is no longer helping you retrieve a chat.
It is helping you regain momentum.
But this is where discovery has to stay disciplined
It would be very easy to stop here and say:
“Great idea. Build it.”
That would defeat the entire purpose of the exercise.
The next question should be:
What would have to be true for this concept to work?
Would people naturally describe old work by meaning?
Would richer result cards actually help them recognise the right work?
Would they trust an AI-generated “where you left off” summary?
Would it reduce rereading?
Would it genuinely make continuing easier than starting over?
Those are testable questions.
So the next step would not be engineering.
It would be a prototype.
Give users a realistic retrieval task.
Observe how they behave today.
Then let them try the concept.
Measure whether they find the right work faster, recover the right information more easily, and feel ready to continue without rereading half the conversation.
Maybe the concept works.
Maybe only one part works.
Maybe users tell us they don’t want summaries at all—they just want better in-chat navigation.
That is the point.
Discovery is allowed to change the answer.
The part of this exercise I found most useful
The most valuable thing I produced was not the solution.
It was the change in framing.
I started with:
ChatGPT has a search problem.
I ended with:
Frequent knowledge workers may have a continuity-of-work problem.
That change happened because I forced myself to stay in the problem long enough before designing the feature.
And that is probably the biggest lesson from this case study.
A stakeholder might give you a feature request.
A dashboard might show a drop.
A customer might complain.
All three are signals.
None of them automatically tells you what to build.
The work in between is product discovery.
For this exercise, I followed a simple sequence:
Verify the problem.
Break it into meaningful segments.
Locate where the journey fails.
Form hypotheses.
Validate what the available evidence supports.
Define the opportunity.
Explore several solutions.
Identify the riskiest assumption.
Test before committing.
It sounds obvious written down.
In practice, the difficult part is resisting the urge to jump from the first symptom to the first plausible feature.
Want the complete case study?
I documented the full exercise in a detailed Product Discovery Report.
It includes the research approach, social-listening evidence, user segmentation, journey analysis, hypotheses, opportunity definition, solution exploration, assumption mapping, research synthesis, prototype test plan, interview guide, and a reusable Product Discovery template.
Download the Full Product Discovery Report →
If you have any trouble accessing the report, or want to discuss the approach, you can reach me at:
Sometimes the most useful output of product discovery isn’t a better solution.
It is realizing that you were solving the wrong problem.