Design and operate scalable LLM pipelines (e.g., Claude, Gemini, GPT-4-class) including prompt engineering, structured outputs, evaluation frameworks, retrieval/vector search, OCR-based document understanding, grounding, guardrails, and hallucination mitigation. Architect distributed, event-driven, streaming services that support sub-second clinician experiences under heavy load, with strict SLOs and strong fault tolerance.