Belal bhai runs a tailoring shop with two machines, in a lane behind Krishi Market in Mohammadpur, and for eleven months of the year he is a calm man. Then Ramzan arrives and the shop turns into a small factory. The stacks of cut cloth climb until they brush the ceiling fan, his assistant starts sleeping under the cutting table for the last ten nights, and the Singer runs past midnight while the mosque across the lane runs taraweeh. Before chand raat he'll have cut punjabis for something like two hundred customers, and no two of those customers have the same body.
That last part matters more than it sounds. A tailor can't run a unit test on a punjabi. He can measure carefully, keep his seams straight, follow a pattern he has cut a thousand times, and he still won't know whether the garment actually fits until one specific human being puts it on at seven in the morning on Eid day. Every one of those two hundred garments is a separate bet, because the pattern is the same but the bodies never are.
So Belal bhai keeps a metric, though he'd laugh at the word. In the weeks after Eid, two kinds of customers walk back into the lane. One kind comes carrying new cloth, because the punjabi fit and now the office needs trousers, or a younger brother is getting married and has been dragged along to be measured. The other kind comes with the Eid punjabi folded in a bag, to have the shoulders taken in or the seams let out. I once asked him how he decides whether a season went well, and he didn't mention money at all. "Eid er por koyjon notun kapor niye ashe, ar koyjon purano ta khulte ashe. Oitai amar hishab." How many come back with new cloth, and how many come back to open the old stitches? That's my accounting.
A product that can't be fully verified before it ships has to be judged by what walks back through the door afterwards, and the only honest way to do that is to count both queues, not just the one holding a complaint.
Classical software never needed Belal bhai's accounting, because it could be tested like a bridge. If pagination worked in staging, it worked for everyone, and QA could sign the release the way an engineer signs a load calculation. Agentic products have quietly taken that certainty away. The same prompt produces different outputs on different days, and the model that behaves beautifully on the eighty cases you thought to grade will do something baffling on case number eighty one. Your eval suite, however honest, only covers the subjects you knew to put on the report card. The customers keep walking in with bodies you never measured. I've argued before that the support inbox is the eval set you didn't write, and this piece is its sequel: if the inbox is an exam you never wrote, the praise-to-criticism ratio is the grade coming back on it.
What most companies reach for instead is NPS, and it's worth remembering what NPS actually is. Fred Reichheld introduced it in a 2003 HBR piece titled The One Number You Need to Grow, and the mechanics are a survey question. How likely are you, on a scale of zero to ten, to recommend us to a friend? Subtract the detractors from the promoters and report the difference to the board. The trouble is that the question is a hypothetical about a future act, asked of the small slice of customers who answer surveys, at a moment the company chose because the customer had just done something pleasant. Later research was not kind to it either. Keiningham and his coauthors went looking for the claimed link to growth and found NPS predicted revenue about as well as the ordinary satisfaction measures it was supposed to replace.
Meanwhile the actual verdicts arrive every day, unasked: support tickets, app-store reviews, replies under the launch tweet, the two-line email from a stranger saying the export feature saved him a day of work. Nobody prompted those people. They found the door to the lane and walked back in on their own, exactly like Belal bhai's two queues. And almost nobody tallies them, because the stream feels too messy to be a metric. The same non-determinism that broke our testing also gave us models that can read ten thousand feedback events and label each one in seconds, which means the messy stream has become classifiable at almost no cost. Label every inbound piece of feedback as praise, criticism, or neither, divide the one count by the other, and plot the ratio against your releases to see which way it bends.
What NPS asks
What the tally counts
Now for the part where I have to be careful, because this idea has a famous corpse in its family tree. In 2005 the psychologist Barbara Fredrickson and a consultant named Marcial Losada published a paper claiming that human flourishing required a positivity ratio above exactly 2.9013, a constant Losada said he had derived from equations in fluid dynamics, applied to transcripts of business meetings. HBR later ran a management version of the claim as the ideal praise-to-criticism ratio, and the figure toured the conference circuit for years. Then in 2013 Nick Brown, a master's student studying part time, teamed up with the physicist Alan Sokal and the psychologist Harris Friedman and took the mathematics apart in public. The Lorenz equations had been applied to meeting transcripts with no justification whatsoever, the constant with its four decimal places was numerology, and Fredrickson ended up formally withdrawing the modelling half of the paper. If you sanctify a universal constant for the ratio, you will end up exactly there, defending 2.9013 against a master's student with a calculator.
But the corpse teaches the lesson rather than killing the idea, because the same literature contains the version that held up. John Gottman didn't ask married couples anything about their satisfaction. He sat them in a room, had them argue about something real, and counted observable behaviours, every sarcastic roll of the eyes and every small touch on the arm. The couples whose ratio of positive to negative moments during conflict sat around five to one were the ones still married years later. He could watch an argument that lasted fifteen minutes and predict divorce with an accuracy that still unsettles people. His method survived the scrutiny that flattened Losada's because he had counted what couples actually did instead of fitting a constant to what they said, and a feedback stream is the same species of evidence, a pile of things customers actually did.
Count events, not sentiment scores; a binary label survives model changes and mood. Keep the classifier prompt versioned so the metric means the same thing in March and in November. Normalize by active usage, because criticism volume grows with adoption even when the product is improving. Segment by surface, since one broken feature can drown ten healthy ones. And never bonus a team on the ratio, or within a quarter you'll have built a Pathao driver.
Praise is structurally underreported, and that asymmetry works in the tally's favour rather than against it. An annoyed customer will find your support form at two in the morning, while a delighted one mostly just uses the product again and tells nobody. Writing an unprompted thank-you costs effort that almost no satisfied person spends, which makes each one that does arrive worth far more than a checked survey box. The bias is real, but it's stable, and a stable bias doesn't corrupt a trend line. If the praise share of your feedback stream is climbing quarter over quarter against that headwind, something in the product is genuinely working, whether or not any survey noticed. I don't know what the right ratio is for a rising product, and after Losada I'm suspicious of anyone who claims four decimal places. But I'm fairly sure the ratio exists, and that ten thousand teams tallying it honestly for a few years would find its shape the way Gottman found his, from the counting rather than from the equations.
Belal bhai has never run a classifier over anything. When I stopped by the lane the Friday after last Eid he was marking a hem with the shutter half up, and he told me three customers had come back that week carrying new cloth, against one punjabi returned in a bag with a sleeve problem he blamed, with some justice, on the customer's new gym membership. He would not call that a ratio. He ordered new thread on the strength of it anyway.