A chat assistant on models we train and serve: next-token framing, three training stages, resumable SSE streams, prefix caching, token quotas, safety, evaluation.
Engineers designing AI products and the systems that serve them.
Your approach: Explain the AI behavior and system boundaries in the brief, with evidence for the requested quality, privacy, latency and cost.
Design an LLM Chat Assistant like ChatGPT. A chat assistant on models we train and serve: next-token framing, three training stages, resumable SSE streams, prefix caching, token quotas, safety, evaluation. Work from the scoping questions below. State assumptions for any unspecified load, guarantee or target, then trace your design end to end. Explain one difficult case and a credible alternative; the worked example is a reference, not a required implementation.
Resolve the scoping questions for an LLM Chat Assistant like ChatGPT. Separate stated behavior from assumptions, and identify what is outside your design.
Declare relevant volume, latency, freshness, quality or cost targets with units. Show calculations or an evaluation plan that can test them; unspecified targets are your assumptions, not hidden pass criteria.
Explain how your guarantees hold in a difficult case relevant to this subject. Address: Search past chats? Browsing, files, memory across chats?
Identify users, required behavior and exclusions. Answer: How many users? Our own models?
Define the information owned by the system and the inputs, outputs and errors at its boundaries. Resolve: How fast? Long chats?
Estimate the dominant workload and resource demand with units and explicit assumptions. For a learned system, also state how quality is measured and what data is available.
Draw or describe the responsibilities needed for an LLM Chat Assistant like ChatGPT. Trace a representative request, event or job from its input to a visible result; identify durable state owners.
Walk through a difficult case step by step, including detection and recovery. Consider: Dropped connections? Limits? Search past chats? Browsing, files, memory across chats?
Compare a credible alternative using your chosen workload and guarantees. Explain a remaining risk, a signal to watch and when you would change the design.