AI in our own server room: when too many requests arrive at once
Our AI models run on our own hardware. When too many requests arrive at the same time, they fail for want of resources. What helped us wasn't more hardware, but a clear order in which they get processed.