AI infrastructure
The three numbers that decide what your inference costs
Cost per thousand requests is an SLI. Batching, utilisation and the tail of your latency distribution set it, and most teams are only watching one of them.
Author
Articles by p10node.
3 articles
Cost per thousand requests is an SLI. Batching, utilisation and the tail of your latency distribution set it, and most teams are only watching one of them.
Most cost programmes start with reserved instances and stall. The savings are real but they are last, not first. Here is the sequence we use and why.
Two weeks inside a production system, and the eight questions that decide the score. Most of them are not about the infrastructure.
Two weeks inside your platform, a scored report, and a roadmap your team could run without us.