129. Design an Image-Text Training Data Pipeline
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns.
Pick a system. Work through the problem. Compare your approach.
Company tags are community-reported. Counts on cards show how many people reported that design.
Billions of image-text pairs from Common Crawl: polite fetching, CLIP scoring, dedup, recaptioning, WebDataset shards and takedowns.
Prompt to four images in seconds: latent diffusion, guidance, a distilled few-step student, recaptioned data, safety on both sides, watermarks and C2PA.
Captions read aloud as alt text: encoder, bridge and decoder, CLIP-filtered data, beam search, hallucination-aware evaluation, lanes on one GPU pool.
Find and act on posts that break the rules in text, images, video and live: hash matching, a calibrated multimodal model, review ranked by expected harm, prevalence: one-hour boards for junior, senior and staff, with the theory.
Find the right videos for a typed query among billions: BM25 and a dual encoder over text, speech and frames, LambdaMART on debiased clicks, human raters.
Short clips from a prompt: latent video diffusion, recaptioned data, spacetime attention, cascades, multi-GPU jobs, previews, fair queues, provenance.