Inference cost vs UX: why cheaper AI models can quietly reduce engagement (Amplitude)
Amplitude ran a real production test swapping an agent’s model to cut cost. Conversion held, but latency doubled and people asked fewer questions. The point: your success metric needs to include ‘time-to-answer’ and ‘messages per session’, not just dollars and a top-line conversion proxy.