While each project involves unique tasks, contributors may: Evaluate clinical EHR vignettes paired with a question, a proposed answer, and a distractor “trap” answer across diagnosis and treatment tasks, spanning cardiovascular, nervous, hematologic, respiratory, digestive, urinary, reproductive, musculoskeletal, and integumentary systems; Score the clinical reasoning quality of benchmark items: vignette accuracy and completeness, whether the vignette gives the answer away, answer correctness and gradeability, and trap quality; Check whether the reasoning chain reaches the answer from vignette facts alone, correct reasoning traces, and write short rationales; Work within three independent blind reads, followed by physician adjudication. Ideally, contributors will have: Medical degree (MD or DO) and an active, unrestricted US medical license (verified with the issuing medical board); Board certification or completed residency training, with recent direct patient care (attending-level preferred); Broad, multi-system diagnostic and treatment experience (internal, family, emergency, or hospital medicine); Strong clinical reasoning: differential diagnosis, next-step management, and application of evidence-based guidelines; Prior experience in medical AI evaluation, clinical content or exam-question review, or medical annotation/QA (a plus); Strong written English (C1+).