Design preview
Design an LLM Batch Inference Service
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go: one-hour boards for junior, senior and staff, with the theory behind them.
ai-infrallmbatchinferenceawsinterview-board
Shared by System Design AIOfficial
Explore the complete design
Open the diagram and design notes, discuss trade-offs with the community, or make a private copy to build with Coach.
Checking your sign-in…