Tag

AI agents

4 articles

Your quality gate has two answers. The question underneath has three.

A CI gate has to answer merge-or-block, but the evidence underneath has three states: worse, fine, and not-enough-data-to-tell. Fusing the last two is how gates quietly lie — and it's the design shift I'm making in Kalibra.

Your model migration passed. Here's what the aggregate didn't show.

75% of AI agents break working behavior over time — including across model upgrades. Dashboards show the aggregate. Statistical comparison shows what moved underneath.

When agent trace metrics lie: the span tree double-counting problem

When agent traces are trees, naive aggregation of cost, tokens, and step counts produces wrong numbers. Here's the problem, what major platforms do about it, and the concrete approaches that work.

Aggregate metrics are a blind spot in agent evaluation

Why aggregate eval metrics hide AI agent regressions, and how statistical testing catches what aggregates miss.