simplex

Blogs · 13 Aug 2026 · 13 min

The document was wrong

Most failed agents die on a dirty help centre. Brief from the tickets, refuse the pages that are not true, and review the memory every month.

Josip Tomić

Partner

Partner at Simplex. He leads the AI agent work on YourGPT — the knowledge, the functions, and the site the agent sits on — and writes the studio reviews.

The agent read the article it was given and reported what it said. The article was a year out of date. That is how most agents fail — not hallucination, obedience. The project is the brief from live tickets, the pages you refuse to ingest, and the review that keeps the memory true.

Aurelia’s first problem was not the model.

It was a help centre that had been written for a different product version, a refund paragraph that contradicted the billing page, and a set of “temporary” workarounds from 2024 that were still live. A previous chat tool had been pointed at all of it. The tool was fluent. The answers were often wrong in the way a confident intern is wrong: specific, sourced, and a year out of date.

The team switched that tool off. We did not start the next attempt by picking a platform. We started with the tickets.

This is the piece of agent work that does not photograph well. A table of what people actually asked, next to the pages that claimed to answer them, next to the pages we would not let an agent see.

Why do most AI agents fail after a clean demo?

Because the demo is run on the five articles someone is proud of, and production is run on everything else.

A help centre accumulates. Marketing publishes a pricing page and never takes it down. Support writes a workaround for a bug that was fixed. Product ships a new cancellation flow and the old article stays. Sales has a one-pager from a previous quarter. The intern’s onboarding doc is still shared.

The usual post-mortem blames hallucination. The usual case is simpler. The cancellation flow moved last month. The article did not. The agent read the document it was given and reported what it said. The document was wrong.

That is the failure mode we see. Not invention. Obedience to a stale page.

If you ingest the whole centre unedited, you have hired a librarian for a room where half the books are from the previous edition.

Salesforce reports that 61 percent of customers would prefer to resolve a simple issue themselves, and that 79 percent of service leaders now say investment in AI agents is essential. Self-service only works if the thing being served is true. An agent pointed at a contradictory centre will give those people a fluent dead end.

The demo will still have gone well.

How do you brief an agent from live tickets?

You start in the queue, not in the CMS.

Aurelia’s brief was written from real tickets. The help centre became memory after the jobs were named. That order is the project. Reverse it and you will spend a month tagging articles nobody asked about, then wonder why the agent still cannot finish an SSO loop.

The method is a sequence, so it is a list.

  1. Export a stretch of tickets and chats. A month is enough to see the shape. Keep the customer’s words. Do not rewrite them into your house style.
  2. Group them by the unit of work. “SSO loop after invite.” “Invoice for a seat we removed.” “Webhook retry after a 401.” Not by sentiment, not by channel, not by “billing-ish.”
  3. For each unit, write the finished job: the answer, the fields, the case, the person if any. This is the same test we use in chatbot vs AI agent. If you cannot name the finished job, you cannot brief an agent.
  4. Open the help centre and look only at the units you just named. What is missing. What is stale. What contradicts something else. What is written for a plan you no longer sell.
  5. Repair those pages, or write the missing ones, before anything is ingested. One article, one job. Exceptions next to the rule, not three screens down.
  6. Only then connect memory. Only then talk about a surface.

Elena Marchetti, customer experience at Aurelia, said the thing we want the implementation to earn:

They treated the agent like a piece of the brand, not a side project. It sounds like us, and the team actually uses it.

It sounds like them because the tickets sounded like them. The repaired centre uses the words the queue already used. An agent trained on product-marketing English will miss the customer who says “it just loops.”

The desk, not the widget
Live tickets, related knowledge, a conversation already in progress.

What should you refuse to ingest?

More than you think. Refusal is the quality bar.

A crawl of the whole site feels like progress. It is how you poison the memory in an afternoon. The agent cannot tell a current policy from a landing page that still ranks.

Refuse these, by default.

Marketing pages that describe a plan, a price, or a feature in the language of a campaign. They go stale on purpose. They are written to convert, not to govern. If a number has to be true, it lives in a dated policy article, not in a hero block.

Old versions that were never unpublished. The 2024 cancellation flow. The retired SKU. The workaround for the Android bug that shipped its fix in March. If a human would need a Slack message to know it is dead, the agent will treat it as live.

Duplicates that disagree. Two refund windows. Three SSO articles. A sales FAQ that says “cancel anytime” next to a contract page that says 14 days. A person knows which one is current. An agent will retrieve whichever chunk scores. In billing, refunds, security, and anything contractual, that is a business risk, not a copy problem.

Internal notes, Slack exports, and ticket dumps with people in them. Names, emails, account IDs, the angry paragraph a customer did not expect to become training data. You can learn from tickets without ingesting the tickets. The learning is the brief. The memory is the cleaned article.

Drafts, “WIP,” and the intern’s side doc. If it is not approved, it is not memory. An agent that quotes a draft will be right until the draft is wrong, which is soon.

Thought-leadership blogs that are not procedures. A post about “how we think about uptime” is not an incident line. A post about “the future of billing” is not the refund rule. Keep the blog for buyers. Keep the centre for the desk.

Anything the live system should answer instead. Order status. Seat count. Whether the invoice was paid. Whether the room is free. Whether the slot is open. Documentation can explain the process. It cannot see the account. An agent that recites the shipping window when the customer asked about their parcel is a chatbot. We drew that line in the inbox piece.

Aurelia’s previous tool had been allowed to see too much. Switching it off was cheaper than arguing with it.

What does a clean article look like, if an agent has to use it?

One job. Current. Pullable in a single chunk.

A page that covers setup, pricing, troubleshooting, and cancellation for two plans is a fine page for a human who can scan. Retrieval does not scan. It lifts a passage. If the passage is the enterprise cancellation rule and the customer is on monthly, the agent now has a correct sentence about the wrong person.

Write the article the way you would brief a careful junior.

  • Title it with the words in the tickets, not the words in the information architecture.
  • State the rule in the first short paragraph.
  • Number the steps if there are steps.
  • Put the exception next to the rule.
  • Separate plans, regions, and product versions into their own pages.
  • Date the page. When the flow moves, the page moves the same day.
  • Say when a person has to take it.

Collect what you have. Delete the outdated and the duplicate. Tag lightly. Always leave a path to a human. Import everything first and clean later is how last year’s workaround gets quoted on a Tuesday.

The loop is the product. The platform is the place the loop runs.

What does the monthly review actually look like?

A morning with the team. A sample of threads. A list of pages to change.

Not a dashboard of “deflection.” Deflection is a vendor number. It will go up if the agent is confidently wrong and the customer gives up. We do not manage to that.

The review is closer to an edit meeting.

We pull a sample: threads the agent finished, threads it handed off, threads a person then had to undo. We read them in the customer’s language. For each: did the correct answer exist, was it current, was it direct, did two pages disagree, did the rule change by plan, should a person have had it.

Then we change the memory. We do not change the model and hope. We unpublish the stale page. We split the mixed one. We write the missing exception the queue already knew. We add the new unit of work that showed up this month and was not in the brief.

Aurelia’s work includes that review on purpose. Service design from live tickets. An agent on their knowledge. Automations into the existing desk. A monthly quality pass with the team. The last item is how the first three stay true.

An agent that is not reviewed is a help centre with a mouth.

The same habit belongs on a night line. Exceptions that stump the overnight agent are tomorrow’s articles, or they will stump it again. That is why nights are covered is a knowledge problem as much as a voice problem.

Where does the automation sit, if knowledge is the project?

Next to the memory, not instead of it.

A clean article can tell the agent how a refund works. It cannot issue the refund, open the case, or tag the product. Those are tools. On Aurelia they run through n8n into the desk the team already had. The agent decides the unit of work from approved knowledge. The workflow writes the record.

If you are choosing a canvas for that layer, the comparison is n8n vs Zapier vs Make after ninety days of real workflows, not after a template gallery. The knowledge project does not care which logo is on the workflow. It cares that the case exists, and that a failed write is visible.

Ecommerce desks hit a sharper version of the same split. The order is live data. The return policy is knowledge. Confuse them and the agent recites a shipping FAQ at a parcel sitting in a depot. We wrote the builder question separately, as best AI agent builder for ecommerce. The knowledge question does not change.

When is it time to name the surface?

After the brief can be read aloud. After the refuse-list is written. After you know which desk the case lands in.

We name YourGPT here, late, because that is where it belongs. The longer product review is YourGPT review: what it does after the demo. It is the agent platform we ship: knowledge, tools, a handoff, the brand on the surface. It is not the project. Pointing it at a dirty centre will produce a fluent mess. Pointing it at a briefed, refused, reviewed memory will produce the desk Aurelia actually uses.

Their own essay on why agents give wrong answers is the line we already used in the room: “The AI didn’t hallucinate. It read the document it was given and reported what it said. The document was wrong.” Their homepage is shorter: “A chatbot answers. An agent finishes the job.” Finishing the job requires the job to be written down. YourGPT will not invent the job for you. Neither will we, from a sitemap.

The agents service is designed that way on purpose. Support, sales, operations. Built with memory, tools, and a handoff the team will trust. Reviewed after launch. Not left as a demo. If someone asks us to “just connect the help centre and go live,” we say no. That sentence is how the last tool got switched off.

Those search facts belong in how to get cited. They are not a reason to stuff a help centre with keywords and call it memory. The article the queue needs has to be true.

What does this mean for a team that already has a centre?

The centre is a product, with an owner, a date, and a review. Someone has to be allowed to unpublish. Someone has to be in the room when product changes a flow. Someone has to read a sample of threads every month and come back with pages. If that person does not exist, the agent will rot on a schedule.

Budget the writing. Quotes that assume “we’ll ingest what you have” are quoting a crawl. The hours are in the tickets, the refuse-list, and the rewrites. The platform licence is the smaller line. Aurelia already had a centre. The work was to make it true enough to speak. The support desk case is that work.

FAQ

Why do AI agents give wrong answers?

Usually because the document they retrieved is stale, contradictory, or mixed with another job. The model reads what you gave it. YourGPT’s own account of a moved cancellation flow is the pattern: the article was wrong, the answer was faithful. Changing the model will not fix a refund window you never updated. Fix the page, or refuse it, then test against the tickets that already proved the hole.

What should I refuse to ingest into an agent?

Marketing pages with prices or plans, unpublished-but-still-live old versions, duplicates that disagree, Slack and ticket dumps with people in them, drafts, thought-leadership that is not a procedure, and anything a live system should answer instead. A crawl of the whole site feels like progress. It is how last year’s workaround gets quoted on a Tuesday. Refusal is the quality bar.

How do you brief an agent from live tickets?

Export a stretch of real tickets, group them by the unit of work, and write the finished job for each unit. Then open the help centre and repair only what those units need. Launch on the jobs you can finish. Tone of voice comes after. A brief that starts in the CMS will ship a fluent librarian for the wrong library.

Can I just crawl the help centre and go live?

No. A crawl copies the rot. If two pages disagree, the agent will pick one and sound sure. If a workaround is still published, the agent will teach it. Go live on a small set of approved, single-job articles that match the queue. Add memory as the review proves the next unit is clean. “Connect everything” is how the last widget got switched off.

How often should agent knowledge be reviewed?

Every month, with the team, against a sample of real threads — and the same day a product flow changes. The monthly pass catches the slow rot. The same-day pass catches the cancellation page you otherwise leave live for three weeks. An agent that is not reviewed is a help centre with a mouth. We treat the review as part of the build, not as an optional extra.

Is YourGPT the project?

No. YourGPT is the surface we use: knowledge, tools, a handoff, the brand on the front. The project is the brief, the refuse-list, and the review that keep the memory true. Point any capable platform at a dirty centre and you will get a fluent mess. We name YourGPT late so nobody confuses a licence with the work.

What does a monthly quality review look like?

A sample of finished threads, handoffs, and reversals, read in the customer’s language. For each, we ask whether the answer existed, was current, was direct, and should have been a person’s. Then we unpublish, split, or write. We do not manage to a deflection percentage. We manage to pages that are still true.

Do we need a new help desk to clean the knowledge?

No. The desk is the system of record. The knowledge is the memory. Automations can write the case into what you already run. A vendor that requires a migration before they will let you refuse a stale article is selling you a new inbox. Keep the desk. Repair the pages. Then connect the agent.

Send the current help centre and a week of tickets. We will write back with what an agent should answer, what it should never touch, and which pages we would refuse to ingest. That note is the start of the agents work, and it is how we prefer to be reached.