Evaluating LLM Workflows: A Field Guide
What mature LLM evaluation actually looks like in 2026, and why squinting at outputs in a spreadsheet doesn't count.
Read essay →I'm Matt Ferrante. For 15+ years I've shipped software in AdTech, algorithmic trading, and now AI. The kind that has to actually work on Monday morning, not just demo on Friday.
Pick the one that matches where your team is. All three can convert into each other.
For teams that pitched an AI project, got the budget, and eight months later have nothing shipped. Get out of pilot purgatory and onto the P&L.
Ship AI-native features, agent APIs, and MCP endpoints that survive real customer traffic. Draws on 18 MCP servers in production, one at 100k requests a day.
For leaders whose 'what's your AI strategy?' answer needs to hold up in a board room, and for teams that need real fluency, not another slide deck.
What mature LLM evaluation actually looks like in 2026, and why squinting at outputs in a spreadsheet doesn't count.
Read essay →A checklist for making a codebase pleasant for coding agents to operate in: file layout, naming, docs.
Read essay →Most software companies think AI is a feature. The real question is whether AI can replace what you sell.
Read essay →Start with a 30-minute consult. $20, refunded if it isn't useful. Bring your worst architecture decision, a stalled rollout, or a demo you can't put in front of customers. You'll leave with a plan.