Saquif

Topic · AI

Writing on agentic AI, RAG, and LLM craft

Essays on agentic platforms, RAG pipelines, evals, prompt design, and what it means to ship AI inside a B2B product without losing your shirt.

13 posts in this topic. See all writing.

A vendor pitch in March 2026 promised “99.7% retrieval accuracy” on a benchmark the vendor had built themselves, against questions the vendor had written, evaluated against a corpus the vendor had cleaned. That meeting is the unstated prior under “The half-life of a benchmark,” “Why your RAG retrieves the wrong thing 14% of the time,” and “Eval suites are the new release notes.” Across eleven pieces the cluster argues that the AI features most enterprise PMs are scoping right now will be obsolete before the next planning cycle, partly because the cost curve and the evaluation surface are moving faster than the roadmap can. “Pricing AI products by token is a trap” sits next to “The compounding cost of LLM context” because the unit economics of agentic workflows look fine in the demo and terrible in the third month of production. “AI eats QA before it eats engineering” is the operational corollary. The reader’s payoff from working through the cluster is a usable mental model for which AI bets compound and which expire on the slide deck, well before your CFO starts asking questions about gross margin in the Q3 review.