simplex

Blogs · 21 Aug 2026 · 16 min

AI agent development services: what you are buying after the demo

What belongs on an AI agent development invoice, how to audit a quote, and what to accept on the Monday after launch — without buying another chatbot.

Josip Tomić

Partner

Leads AI agent work and innovation at Simplex — the knowledge, the functions, and the site the agent sits on — and writes the studio reviews.

AI agent development services are the paid work of briefing an agent from last week's tickets, refusing knowledge it must not use, wiring actions into the desk you already run, and writing a Monday test you can withhold the last invoice against. A quote that names a framework and a demo, and never names the inbox, is not a service.

The SOW arrived on a Thursday. Twelve pages. An orchestration framework, an eval harness, a six-week proof of concept, a slide about multi-agent. No mention of the help desk they already pay for. No mention of Monday.

That packet shows up on most first emails. A week of tickets shows up on the honest ones.

The artefact on most quotes is a demo. The work that survives is last week's inbox as the brief, a list of things the agent will not touch, actions in the desk you already run, and a Monday you can withhold payment against.

If the quote never names the inbox, it is not a service.

The definition of the work is in chatbot vs AI agent. How we read an export is in the tickets already wrote the brief. This page is the commercial one: what you are buying, and what to accept when the demo is over.

What line items belong on an AI agent development invoice?

Not a framework. Six jobs. If a quote cannot name them, you are buying a mood.

Brief. A paid line for last week's groups, in the customer's words — not a free discovery workshop. If that line is missing, they are briefing from a deck.

Refuse-list. A paid line that names pages and tools the agent will not touch. If that line is missing, they will ingest the centre and call it memory. The edit work is the knowledge project. The invoice test is whether they will write the nos down.

Actions in the desk you already run. Open the case. Tag the product. Attach the ID. Write the night note. If the work cannot change a record, you have purchased a conversation. Keep the desk. Do not migrate it as a precondition.

Handoff that carries the thread. Transcript, fields, one-line reason. A customer who has to repeat themselves was dropped. Zendesk's CX Trends 2026 (November 2025): 85 percent of CX leaders say customers will drop a brand over unresolved issues — even on the first contact.

Surface that is not the project. Widget, WhatsApp, voice. The surface wears the brand. It is not the invoice.

Review after launch. A sample of threads, every month. What the agent finished. What it filed and then got wrong. What new unit showed up. A build with no review line is a demo with a longer calendar.

LineWhat you can point atWhat a widget installer writes instead
BriefLast week's groups, named“Discovery workshop”
Refuse-listPages and tools marked no“We’ll ingest the sitemap”
ActionsThe desk, the fields, the log“Omnichannel experience”
HandoffA complete thread in the queue“Talk to a human” button
SurfaceOne place it livesThe project itself
ReviewA date after Monday“Hypercare, two weeks”

The invoice is the six jobs. Everything else is decoration.

Salesforce’s 7th State of Service (April–June 2025) found 79 percent of service leaders call AI-agent investment essential, while representatives spend 46 percent of their time with customers. An invoice that cannot name which units leave the other 54 percent is buying the headline.

Should you hire a studio, buy a platform, or staff engineers?

At Think on 5 May 2026, IBM CEO Arvind Krishna said: “The enterprises pulling ahead are not deploying more AI; they’re redesigning how their business operates.” That is the hire test. You are hiring for the part of that redesign you cannot do yourself.

Hire a studioBuy a platformStaff engineers
The job you cannot doBrief, refuse-list, Tuesday ownershipOne named job, already true in the centreA product surface you will extend for years
You already haveA desk, a week of tickets, someone who can unpublishA clean centre and a person who will sit with itA desk owner and engineers who will keep it
You do not haveHours to learn the platform on productionThe hours to brief and reviewA reason the platform cannot file the case
FailureStudio ships a widget and leavesPlatform sits unused after the trialCustom stack nobody on staff can operate

Buy the platform when one job is obvious, the articles that answer it are true, and someone on your team will read a sample next month. A trial is enough to fail fast. It is not enough to skip the brief.

Hire a studio when the invoice items are the work, and nobody on staff owns them. That is the agents service we sell. Aurelia was this shape. So was Vesper on the night line.

Staff engineers — or an engineering firm — when the agent is a product you will extend for years, the desk already has an owner, and a platform cannot file the case. LangGraph is a software programme, not a first agent. On this query the named shop is usually N-iX.

Constraint is the existing desk and who owns Tuesday, not “AI maturity.”

Anil Jain, in Google Cloud’s December 2025 recap from more than 3,466 executives: “The era of scripted chatbots and reactive customer service is coming to an end.” If the plan is still building chatbots to answer questions, you are buying the last decade. Hire against the desk.

When should you hire N-iX for AI agent development?

When the agent is a software programme you will extend for years, and a platform cannot express the job.

N-iX sells custom agents, multi-agent orchestration, and integration into CRMs, ERPs, and on-prem stacks. Their page names LangChain and LangGraph as the production frameworks, an APEX sequence — Assess, Pilot, Expand, eXcel — and a team they put at 200+ AI and ML engineers, 60+ AI and data projects, and 2,400+ technologists. They say standard implementations take a few months, with a proof of concept before scale. Those are their figures, on that page, as of 2026.

That is a real buy. It is not the buy for a first support agent on a help desk you already run. Gartner’s Anushree Verma, in the 25 June 2025 note, wrote that “many use cases positioned as agentic today don’t require agentic implementations.”

N-iXA studio (this one)A platform trial
What you are buyingCustom architecture, integration, a team that will keep the codeBrief, refuse-list, actions in the existing desk, a Monday clauseA configured agent you will sit with
Fits whenThe agent is a product inside ERP or a proprietary stackNobody on staff owns the six invoice linesOne job, a true centre, an owner
Time they describeMonths, plus a PoCA week of tickets, then the deskDays, then you keep or kill it
FailureA sandbox that never files a live caseA widget with better copyAn unused trial

Hire N-iX when you already have a desk owner who can unpublish, engineers who will inherit the repo, and a reason LangGraph has to exist. Their “proof before scale” line is the right instinct. Ask them which live desk the proof writes into.

Hire a studio when the work is last week's five groups and a Monday clause. We are the wrong firm for a multi-agent programme across three ERPs. Hire neither when the week is empty — that is not an engineering gap.

N-iX is the engineer pole. A studio is the invoice pole. Do not buy the first for a job that is the second.

How do you audit an agent quote when every page invents a price band?

You do not need our number. You need a way to kill theirs.

The search results for this query are price tables, and the bands disagree. We will not add another one. Gartner, reported by CX Dive on 17 August 2026, analysed 432 customer-service AI use cases: one quarter produced a return, one quarter a negative return, 42 percent had unclear ROI, 11 percent broke even. Teams were running nearly five use cases on about 13 percent of the functional budget. More than three-quarters of leaders still planned to increase spend.

Money is not the scarce fact. A quote you can audit is.

Score the SOW against the six line items. If “discovery” never asks for last week’s export, it is a workshop. If “build” never names the desk, it is a widget. If “launch” has no Monday clause, it is a demo with a go-live date.

Score the cost story against the work, not the tokens. Computerworld, reporting Gartner’s 17 August 2026 note, wrote that token costs will fall while inference costs for agentic workflows rise more than fivefold over two years. Agent inference already costs five times a chatbot; the same task can take five to thirty times the tokens. Gartner’s Will Sommer and Sabine Zimmerhansl call it the inference paradox: buyers assume cheaper tokens become cheaper agents. “They will not.” A quote that sells a reasoning model on every FAQ is buying the expensive path on purpose.

Score the use-case list against Verma. Gartner estimated only about 130 of the thousands of “agentic” vendors were real, and named the rest agent washing — assistants, RPA, and chatbots rebadged. If the quote offers HR, sales, and support in the same first month, it is a menu. Menus do not survive Tuesday.

Julie Geller, principal research director at Info-Tech, told CX Dive: “Too many AI rollouts begin with pressure to demonstrate a credible AI strategy to the board, rather than with a clearly defined business problem.” Containment, she said, is mistaken for success. “Delaying contact with a human agent is not the same as resolving the customer’s problem.”

If the quote saysAsk thisKill it if
$X for a six-week PoCWhich live desk does it write into in week three?The PoC never leaves a sandbox
Eval harness / LangSmithWhat is the acceptance sentence on Monday?Quality is a dashboard, not a queue
Multi-agent from day oneWhich single job failed first?They cannot name one job
“Ingest your site”What will you refuse?The refuse-list is empty
Deflection as the KPIWhat did the customer get?Success is “they stopped writing”

A price without a Monday clause is not a price. It is a hope.

How do you accept the work on the Monday after launch?

Not a satisfaction score from three people who finished a survey. Not a containment percentage. The same five groups from last week, in the desk you already run.

Repeats are already filed. Subject, product tag, ID, the article that was true. A person does not rebuild them from a chat export.

Exceptions arrive with context. The thread, the fields, the reason it left the brief. Nobody is paging through a widget log at 08

.

The refuse-list held. Billing disputes, security, legal, anyone who already ran the documented checks — those did not get a fluent paragraph. They got a person.

A log exists. What the agent finished. What it handed off. What it was not allowed to touch. If you cannot open that list on Monday, you cannot accept the work.

Zendesk’s same report: 74 percent of consumers now expect service twenty-four hours a day because of AI, and 95 percent want an explanation for AI-made decisions. Monday checks both. The repeats were covered overnight. The exceptions can be explained.

Which statements of work should you refuse to sign?

Refuse the packet that opens with the model, asks for a sitemap instead of last week's export, or sells support plus sales plus HR in week one. A first agent is five groups, not three departments. Refuse a kickoff that is a design file for the chat bubble. Aurelia had already switched a polite widget off for that.

Refuse multi-agent as the opening move. Sommer and Zimmerhansl’s note: swarms incur a “massive inference tax” before a user sees a result. Your first job needs a case in the desk.

A proof of concept that never touches the live desk. Sandbox tickets, a recorded demo. Gartner’s June 2025 note said over 40 percent of agentic AI projects will be canceled by the end of 2027 — escalating costs, unclear business value, inadequate risk controls. A PoC that cannot file a real case is how that number is made.

An SOW that owns the prompts and not the knowledge. You inherit a tone and a centre that still lies. The next vendor will blame the model.

Unlimited use cases. That is how the refuse-list dies, and how the inference bill becomes the project.

Replace the desk as a precondition. You are buying a new inbox. Aurelia kept theirs.

Success defined as containment. If the only number is “deflection,” the customer can still be stuck.

No named owner after they leave. Who unpublishes. Who reads the sample. Who owns Tuesday. Missing names mean they are renting you a demo.

We will not sign the inverse either. A one-page “install the agent” with no brief, no refuse-list, and no Monday.

If they cannot tell you what they will refuse, they will refuse nothing.

Who should still own the brief, the refuse-list, and the desk after they leave?

You. If that sentence makes the room uncomfortable, do not hire yet.

The studio can write the first brief. They cannot sit in your queue in November. Someone on your side has to unpublish, mark a new exception, and read a sample without us in the room.

Krishna’s line is the ownership test, not a transformation slide. If you are not redesigning who owns Tuesday, you are deploying more AI. IBM’s IBV 2025 CEO Study (2,000 CEOs, published 6 May 2025) found only 25 percent of AI initiatives delivered expected ROI, and only 16 percent had scaled across the enterprise. The gap they name is readiness, not another model.

Name three people before kickoff.

The brief owner. Updates the five groups when a new one takes the morning.

The refuse owner. Unpublishes. Adds the next no. Does not wait for a vendor standup.

The desk owner. The queue is still theirs. The agent writes into it. They decide when a handoff is a drop.

If those three are the same person, that is fine for a first agent. If they are nobody, the studio will own the project until the contract ends, and then the widget will rot.

We stay on the monthly sample. We do not stay as the only people who can change a rule. A service that cannot leave you with those three names has not finished.

The platform we use for customer and internal agents is YourGPT. Knowledge, functions, a handoff — not a FAQ skin. The product review is what it does after the demo. It shows up late on this page because it is not the hire. The hire is the six line items, the Monday clause, and the three names. If a different platform already files the case in your desk, keep it. We will say so.

FAQ

How much do AI agent development services cost?

There is no honest single band — vendor pages invent ranges that disagree with each other by a wide margin. What you can price is the six line items, the first month of review, and the inference path. A reasoning model on every FAQ is the expensive one. Gartner’s August 2026 work, via Computerworld, says agentic inference is already five times a chatbot and rising more than fivefold through 2028. Ask what the number is paying for. Refuse a quote that cannot say.

Do I need a custom LangGraph build for a first agent?

Not if the job is five groups in a desk you already run. Custom orchestration is for a product you will extend for years, with engineers who will keep it. That is the N-iX shape: LangGraph, ERP integration, a few months, a PoC first. Verma’s June 2025 note still applies: many use cases sold as agentic do not need to be. Start with the brief and the desk. Staff a framework when the platform cannot file the case, and not before.

Is N-iX the right AI agent development company for a first support agent?

Usually not. N-iX is an enterprise engineering firm — custom agents, multi-agent systems, CRM and ERP integration, a team they size in the hundreds. Hire them when the agent is a product you will extend for years and a platform cannot file the case. A first support agent on a desk you already run needs last week's tickets, a refuse-list, and a Monday clause. That is studio work, or a platform trial. Do not buy 2,400 engineers for five repeating questions.

Is a seven-day platform trial enough?

A trial is enough to fail a bad job. It is not enough to skip the brief, the refuse-list, or the Monday clause. Five proud articles will go well in a week. Production runs on everything else. Buy the trial when someone on your team will sit with it. Hire a studio when that sitting is the work.

What should I send before anyone quotes?

Send last week's tickets — or the overnight log if nights are the pain — plus the current help centre, the desk you already pay for, and three things you would rather the agent never touch. Do not send a chatbot wishlist. If a studio quotes without asking for that export, they are quoting a widget. The note we write back from that send is yours to take elsewhere.

Do I have to replace the help desk?

No — keep the help desk you already run and make the agent write into it. The desk is the system of record: cases, tags, IDs, the night note. A migration sold as a precondition is a vendor preference, not a requirement of the work. If the platform cannot file a case where the team already sits, you are buying a second inbox and a reconstruction job every morning. Hire against the existing queue, or do not hire.

When should I not hire AI agent development services?

Do not hire when the week is empty, the centre is a pile of PDFs nobody will unpublish, you want a chatbot for the board and will not accept a Monday clause, or the agent is a product and you already have engineers plus a desk owner. An empty week is a findability problem, not a model problem. Hire against a named gap. Do not hire against a slide.

What happens in the first month after launch?

You accept the first month against the same five ticket groups, in the existing desk, on the first Monday. A sample of threads shows what the agent finished, what it handed off, and what it was not allowed to touch. Pages that proved false get refused or rewritten. A new unit that took the morning goes into the brief or it waits.


Send last week's tickets — or the overnight log — and the current help centre. We will write back within 24 hours with the units an agent should finish, the ones it should never touch, and whether the quotes you already have are a service or a widget. The note is yours. Contact the studio with the export. That is how the agents work starts.