18. Keep enterprise AI search fresh and private
Handle document changes, ACL revocation and deletion across a retrieval pipeline.
Start with a template. Work through each step. Ask Coach when you need a second opinion.
Company tags are community-reported. Counts on cards show how many people reported that design.
Handle document changes, ACL revocation and deletion across a retrieval pipeline.
A chat assistant on models we train and serve: next-token framing, three training stages, resumable SSE streams, prefix caching, token quotas, safety, evaluation.
Open-weight LLMs from 8B to 405B served to many products over a streaming OpenAI-style API on AWS GPUs.
Millions of LLM prompts in JSONL files, one result per custom_id within 24 hours at about half the online price, on GPUs that come and go.
One API in front of every model: token quotas, routing and fallbacks, caching, cost attribution, masked logs and guardrails, streamed without buffering.
Serve 10,000 LoRA adapters for 2,000 tenants on one shared base model, with mixed-adapter batches, tiered adapter caches, affinity routing and base upgrades.
Exact, semantic and prefix caching in front of LLMs: scoped keys, distances with an error budget, tenant isolation, invalidation and measured savings.
Answers from private documents with citations: chunking, BM25 + vector retrieval, reranking, permission filters, grounding checks and evaluation.
Next words within 100 ms of a keystroke: conditional language modelling, a distilled student on draft-pinned KV caches, private n-grams, DP and canary audits.
Translation for 20 to 240 languages: one multilingual encoder-decoder, mined and back-translated data, distilled students, COMET and MQM, quality-aware serving: one-hour boards for junior, senior and staff, with the theory.
Prompt to four images in seconds: latent diffusion, guidance, a distilled few-step student, recaptioned data, safety on both sides, watermarks and C2PA.
Captions read aloud as alt text: encoder, bridge and decoder, CLIP-filtered data, beam search, hallucination-aware evaluation, lanes on one GPU pool.
Faces of people who do not exist: a style-based GAN, truncation, latent-space edits, selfie inversion, provenance and memorisation audits.
Selfies in, professional headshots out: per-order LoRA on SDXL, an identity-encoder preview, face-match ranking, GPU pools, consent and deletion.
Images at 2K–4K without paying for every pixel: latent diffusion, an SR cascade with noise augmentation, tiled decoding, distilled drafts, cost per megapixel: one-hour boards for junior, senior and staff, with the theory.
Short clips from a prompt: latent video diffusion, recaptioned data, spacetime attention, cascades, multi-GPU jobs, previews, fair queues, provenance.