
your robots have amnesia
I told you your context was the problem.
I meant it (that was in April, and I'd write it again today). A real CLAUDE.md is the difference between generic slop and something that sounds like you.
But my context was clean this week and my agents lied to me three times anyway.
Not hallucinations but something quieter, and worse.
Here's the failure nobody's checking for 👇
The failure nobody's checking for
Everybody watches for hallucination. Fair. It's the one with a name, a benchmark, and a thousand boring LinkedIn posts and shitty lead magnets.
But AI has memory problems, it’s causing amnesia to your workflows, and it’s a much bigger problem to have a core job automated.
Mark Hendrickson (he builds Neotoma, the memory layer I run under my own agents) draws the line better than I can:
Hallucination is a model-level failure. The model makes something up.
Memory corruption is an infrastructure-level failure. The stored data is wrong. The model retrieves it faithfully.
Basically, this is what it means:
Your agent isn't lying to you. It's remembering wrong.
But the part that should bother you: memory corruption passes every check built for hallucination.
The citation is real + the source exists. So getting the data is accurate AND every guardrail you own says green 🟢
Well, the fact is still just plain wrong.
So "there's no bad agents, there's only bad context" is true, and it's also where I stopped thinking.
Good context isn't the finish line. Correct context is, and those are different problems with different fixes.
Three times, this week, on my own systems
Because of that amnesia problem, I wasted 10 hours into agents troubleshooting that could have been fixed in 1.
Here’s three moments where my own systems failed this week:
Bad context. I wrote a skill and pointed it at a job. It refused, confidently, for a reason that was flat wrong. I'd baked in an assumption I didn't know I'd made, and it enforced it perfectly.
→ Why you care: your agent inherits every assumption you never said out loud.
→ The fix: when it refuses, ask what rule it thinks it's applying. The rule is usually yours.
No paper trail. An agent flagged a real client quote as fabricated. It was right to. I'd said that quote out loud, and speech doesn't land in a file.
→ Why you care: a true thing with one source and no trail looks exactly like an invented one.
→ The fix: write down where a fact came from at the moment it arrives, not later.
Its own echo. An agent told me a name with total confidence. Thirteen sources for, zero against. It was wrong. All thirteen traced back to one speech-to-text engine guessing at a spoken word, and every file after that copied it (I'd met the person. The machine hadn't.).
→ Why you care: volume is not corroboration.
→ The fix: count sources, not mentions.
None of those is the model being stupid. because very one is a fact that got stored wrong and read back honestly, kinda like a human mistake.
The 20K feet view
A developer has scaffolding for exactly this problem with version control, comments, a senior dev who remembers why the thing was built that way (and hold everything together).
I'm a PMM and I had none of it 😅
So I built the thing I actually needed: one map of the whole system, living outside the build, that I can open cold three weeks later.

It has altitudes:
Map (20,000 ft). One view of the whole system. What each piece does, and what the client is actually paying for.
Agents (10,000 ft). The handful of jobs. Each one is an input, a process, an output.
Workflows (5,000 ft). The sequences that chain the agents together. Where one agent's output becomes another's input.
Skills (1,000 ft). The named moves you trigger by hand. /battle-card, /positioning-audit, /launch.
Tasks (ground). The atomic steps. Read the source. Check the memory. Draft it. Log it.
Something breaks, you start at the top and drop one altitude at a time. You don't reopen the whole thing and pray (three days on what should have taken 1 hour, and I'd built the thing myself).
The map is also how you trace a wrong fact back to where it got in. Corruption is invisible without one, because every layer is faithfully passing along whatever the layer above it said.
Why this hits PMMs harder
You're about to own these systems, and most of you didn't ask to (this is just the way AI is becoming a crucial tool in our toolboxes)
The engineers teaching you to build have twenty years of muscle memory for exactly this problem. Version control isn't a tool they recommend, it's a thing they'd no more skip than breathing. You have a folder, a good idea, and a Tuesday.
And the tooling is generous now, which is the trap. You can generate more system in an afternoon than you could read in a week. More code, more surface, more places for a wrong fact to sit (and none of it written by you, so you can't spot it by reading).
That's the vibe-coding bill. The issue is not the code you can't write but so much code that you can't audit it.
I'm not writing this from the other side of it, by the way. The month is half gone and my own pipeline has nothing new in it, because I spent the week fixing something instead of selling anything 😬
That's what a system without a map actually costs. Not the three days. The three days plus everything queued behind them.
If you'd rather not learn it the way I did, this is what I trained a consultant to do this week.
Two sessions, four hours, live, on your real work:

Session 1 is the Brain: your judgment, your format, your standards, out of your head and into a file, so the agent stops guessing.
Session 2 is the Build: the agent that reads it, pointed at one PMM function you already own. Compete, Position, or Launch. Your pick, and they get harder in that order.
You don't watch me build. You build, and I sit next to you.
What you leave with:
The agent, running on your machine, on your product
The agent map, written during the build (the part people thank me for six weeks later)
A cheat sheet of every move we made, so you can run it again without me
The recordings and the full transcript
A GitHub repo, scaffolded the way a PMM thinks, not the way an engineer does
14 days of Slack while you build the version two yourself
If this sounds interesting, you can get more info on my training here.
I've spent 75 episodes of my podcast interviewing product marketing experts, and worked with 45+ startups. Most of them were rebuilding the same thinking from scratch every quarter.
It comes out of L&D, not your budget. And the point isn't that I built you an agent. It's that on Monday you're the person on your team who builds them (and this is priceless to create influence)
You own these systems now. Accountability comes with that, and accountability without a map is just stress.
You're the operator of the system, not the last person who touched it.
That's why we need to build with AI, not just use it.
Talk soon 👋
- Gab
