case file · PRODUCT · 4:30 pm

chittagong port · gate 3

Subject. Where the failure catalogue is actually written.

Your support team is the eval set you're not using.

In agentic-AI products, the people answering tickets and walking new customers through their first ninety days already hold the failure-mode catalogue your eval team is trying to write from scratch. The under-priced research function is sitting one floor down, badged in a different colour, and the company keeps treating it like a cost line.

open the ledger

Bangladesh Customs · Bill of Entry

CPA · CTG · Day Book · entry no. 04417

HS 8528.72 .. monitors, 40'ft .... gate 3
HS 6109.10 .. cotton t-shirt .. EXEMPT 2
HS 3304.99 .. cosmetics ..... HOLD inspect
HS 7308.30 .. doors, steel .... circular pend.
insp. Hossain
won't sign w/o
original BoL
this week

There's a customs broker who works out of an office on the second floor of a low building near the Chittagong port, two streets back from the container yards, and he keeps a hardback ledger on his desk that he has been writing in since roughly 1996. The ledger has the cardboard cover of a school exercise book, taped along the spine three times over, with each entry written in a small slanted hand and stamped at the end of every shipment line with the broker's seal in faded purple ink. The pages catalogue, in the margins, the things the printed customs forms do not have a column for. The HS codes the deputy commissioner is going to call in for inspection this week sit on one page, and the next page has the inspector at Gate 3 who will not sign the form until you have got the original bill of lading rather than the buyer's photocopy. Two product categories under the new March circular have not yet been updated for in the forms, and the broker has flagged them with a small purple cross so the junior brokers know to ring upstairs before submitting. The shipping companies he files for have been operating, for two decades, on top of this margin commentary without paying for it.

What that ledger amounts to, if anyone in the shipping firms upstream from him ever asked, is a working eval set for the entire customs interface, written by someone who runs the interface forty hours a week and who has never been paid for the formal version of what he knows.

The argument I have been making, half to myself and half to the founders who corner me after talks, is that this is the shape of the missed research function inside agentic AI products in 2026. A new client comes onto the platform without having used a model of this kind before. They do not know how to phrase a prompt that will not derail at the fourth turn, and they do not know which integration their old vendor was handling for them silently before the migration. The customer gets assigned a sales engineer and an onboarding specialist for the first ninety days, and one of those two people watches the platform break in week one in three ways the product team has never seen. The data goes into a Slack channel called something like #onboarding-friction and dies there.

The implementation architect on the third such deployment knows this. By the third, she can sit in the first sales call with a new prospect, hear the customer describe their existing stack, and tell you in the call which two integrations the customer is going to fight in week three. I wrote about this elsewhere in Have we confused thinking with programming?, framed as the deciding skill that has lived outside the engineering room for thirty years. The frame here is narrower and more concrete. The implementation architect's failure mode catalogue, the thing she carries in her head and updates each Tuesday, is exactly the artefact eval teams have been trying to manufacture from PRDs for the last eighteen months. It already exists inside the company doing the asking, sitting on a different desk in a different cost centre.

ticket#ENT-0931
customerLogistics SaaS, Series C
openedday 4 of onboarding
statusresolved · 11h

By the time the new buyer's IT team has finished arguing about the SSO config, I already know which of their two ERP connectors is going to fail in week three. The model can write the patch faster than I can, but the part where you point it at the right line in the connector log is still mine.

quoted, with permission, from the implementation architect's handover note

Hamel Husain has been writing about this for over a year now in a different vocabulary, mostly under the banner of eval driven development and the related idea that production logs and support tickets are the cheapest, highest signal training data a team can get its hands on. The thing his posts keep returning to is that the eval set should be sourced from the place the model actually fails, not from the place the team thinks it might, and the operational consequence is organisational rather than technical. The place the model actually fails is in the seat of the support agent answering the ticket, and the company has to be structured so that what the agent learns is ingested every week, or else the failure stays in a ticket queue and the next ten customers hit the same wall.

This is the same merger that happened to QA between 2024 and the end of 2025, and I argued in AI eats QA before it eats engineering that the eval design skill is now the test engineer skill in different syntax. The next merger, the one nobody has named yet, is QA plus support plus onboarding into the eval function. The senior CSM who has closed seven hundred tickets across three years is not a cost line so much as a research function in a costume nobody updated, doing the labelling work for free and watching the engineering org rebuild a worse version of her instinct from scratch each quarter. The painful part of the merger is not that any of the work is going away but that the org chart has not caught up with where the work has already gone, which means the people doing it are still being paid as if they were doing something else.

The fix, the one I keep recommending to founders who will sit still for forty minutes, is operational rather than philosophical. Run the support queue through a labelling step before tickets resolve, so that every closed ticket leaves behind an eval case, one line long. The label should carry the customer's environment, how the model actually behaved against what was expected, and the line in the log that named the failure. The support lead becomes, by Friday afternoon, the person whose week of work has produced thirty new eval entries the rubric did not previously test against. Pay them on those entries, not just on time to resolve. The frame for hiring this person is already in Hiring for taste. The agentic product version of taste is the muscle that points at the failure before the model has been told to look for it, and the people who have built that muscle have done it on Zendesk and Intercom, not in Cursor.

The broker in Chittagong will write the day's last margin note in the small steady hand he has used since 1996, and the shipping companies he files for will not pay him for the ledger this year either. They will pay him for the filings. The ledger is the thing that makes the filings possible, and it has never once appeared on an invoice, because it is not a deliverable and nobody has ever had to buy it separately.

Every agentic vendor I keep talking to is on that same trajectory, and the shape is close enough to be uncomfortable. The bell on the support manager's desk is a Slack notification. The ledger is a CSV that nobody has imported into the eval suite. And the CSM who is closing her seven hundredth ticket this year is compensated as though the work ended at resolution, when resolution is roughly the point at which the valuable part became possible and everybody walked away from it.