Discussion about this post

User's avatar
Gal Dayan: The Agent Whisperer's avatar

Congrats on the first thousand, Sandipan - well earned. “Beyond the hype” is the whole game. The way I keep seeing it: most AI that demos well is a read system, where being wrong is cheap and you just rerun it. The systems that truly work are the ones that have to act - and the moment an agent acts in the real world, a wrong move costs something you cannot rerun. That gap is where most pilots quietly die. Here’s to the next thousand.

Omri Ben-Shoham's avatar

Congrats on 1,000 - and the thousand you describe sounds like the right one to have. The line that landed for me is the boring part after the demo, because that is exactly where we spend most of our time too, and almost nobody writes about it honestly.

One thing I would add from the calling side of that same problem: the boring part is not just evaluation and data debt, it is also that a lot of production AI now takes actions in the physical world - a call placed, a message sent - and that changes what testing even means. You cannot diff two live phone calls the way you diff two API responses. Curious whether you are seeing that shift show up in the enterprise agents you write about, or if it is still mostly a text-and-retrieval problem for most of your readers.

No posts

Ready for more?