Linguist (Linguistic Engineering & LLM Data Quality)
Location: Fully Remote (US Only)
Contract Duration: 12 months
Pay Rate: $53 - $55/hr on W2
Job Overview
We are looking for a skilled Linguist III to join our team and help build, maintain, and analyze the datasets that power next-generation LLM-driven features for smart glasses and wearable technology.
In this role, you will play a critical part in building and assessing the quality of data processed by production LLM systems. You will work at the intersection of linguistics, data annotation, experimentation, and AI quality evaluation, partnering closely with engineers, research scientists, project managers, and data scientists.
You will help translate complex cross-functional requirements into high-quality datasets, develop and improve annotation workflows, analyze model and rater performance, and provide actionable insights that improve the accuracy, reliability, and overall quality of AI-powered experiences.
This is an opportunity to apply your linguistic expertise to rapidly evolving AI systems while expanding your technical capabilities in Python, SQL, statistics, data analysis, and LLM evaluation.
Responsibilities
- Aggregate requests from cross-functional partners and translate them into actionable, high-quality datasets using synthetic data collections and samples from live user traffic.
- Build, curate, and maintain datasets used for LLM training, evaluation, and quality measurement.
- Develop and maintain manual and automated data-quality processes across multiple concurrent projects.
- Create and optimize LLM quality-grading queues and component-level error-attribution queues.
- Define grading rubrics, write annotation guidelines, onboard and calibrate raters, and continuously improve queue design.
- Design and conduct experiments to evaluate rater quality, guideline clarity, inter-rater reliability, and model performance.
- Analyze numerical rating results and identify meaningful trends, patterns, and quality issues.
- Conduct daily quality audits of annotated datasets, identify inconsistencies, and investigate sources of annotation and model-quality issues.
- Attribute errors to specific LLM components and dimensions, including factuality, brevity, coherence, and related quality metrics.
- Develop datasets and author guidelines rapidly in response to new and emerging LLM data-rating, creation, and annotation requirements.
- Use SQL and Python to extract data, calculate metrics, analyze agreement scores, and support reporting and automation.
- Analyze system-level and subcomponent performance metrics and translate findings into clear recommendations.
- Prepare summary reports and communicate results, milestones, risks, and recommendations to cross-functional stakeholders.
- Collaborate closely with engineering, research, product, project management, and data science teams to align on data priorities and inform product decisions.
- Operate effectively in a fast-paced environment where LLM capabilities, evaluation requirements, and priorities evolve quickly.
Key Projects & Day-to-Day Work
Dataset Building & Curation: Receive requests from engineers and research scientists, then design and execute synthetic data collections or sample existing data to produce clean, structured, and labeled datasets for LLM training and evaluation.
Annotation Queue Management: Stand up and maintain grading queues by defining rubrics, developing annotation guidelines, onboarding raters, and iterating on queue design as new AI capabilities emerge.
Quality Auditing & Error Attribution: Review annotated data, identify inconsistencies, conduct inter-rater reliability checks, and trace quality issues to specific LLM components or performance dimensions.
Data Analysis & Reporting: Use SQL and Python to retrieve and analyze metrics, calculate agreement scores, identify trends, and support dashboards and recurring reporting. Leverage AI-assisted tools where appropriate to accelerate query and scripting workflows.
Experiment Design: Plan and execute controlled experiments focused on rater calibration, guideline effectiveness, annotation quality, and model performance. Summarize results and present recommendations to cross-functional stakeholders.
Guideline Authoring: Draft, test, and refine annotation guidelines for new LLM capabilities on tight timelines, ensuring that raters can consistently interpret and apply evaluation criteria at scale.
Cross-Functional Collaboration: Participate in regular discussions with engineering, research, product, project management, and data science teams to align on data priorities, surface quality insights, and influence roadmap decisions for AI-powered wearable experiences.
Must Have Qualifications
- Bachelor's degree in Linguistics, Computational Linguistics, Speech Science, or a related field, or a graduate degree in Linguistics, or equivalent industry experience.
- 4–6 years of relevant professional experience in linguistics, data annotation, language data, LLM evaluation, computational linguistics, or a related field.
- Experience with data annotation, labeling, or data quality workflows.
- Working knowledge of SQL and Python, with a strong interest in developing deeper technical skills.
- Foundational knowledge of linguistics, including areas such as phonetics, syntax, semantics, or dialectology.
- Ability to analyze numerical ratings and draw meaningful conclusions from data.
- Basic knowledge of statistics and mathematical concepts used in data analysis.
- Strong written and verbal communication skills, including the ability to present findings and communicate project milestones clearly.
- Ability to manage multiple concurrent projects and operate effectively in a rapidly changing environment.
Preferred Qualifications
- Master's degree or higher in Computational Linguistics, Linguistics, Speech Science, or a related discipline.
- Experience evaluating large language models or generative AI systems.
- Experience with LLM data creation, annotation, evaluation, or quality measurement.
- Intermediate Python and SQL skills.
- Experience with statistics, data science, experimentation, or quantitative analysis.
- Experience designing or evaluating annotation guidelines and grading rubrics.
- Experience measuring inter-rater reliability, rater quality, or annotation consistency.
- Familiarity with LLM quality dimensions such as factuality, coherence, brevity, and component-level error attribution.
- Experience using online learning resources such as Coursera, Udemy, or similar platforms to expand technical skills.
Pursuant to the California Fair Chance Act, Los Angeles County Fair Chance Ordinance for Employers, Los Angeles Fair Chance Initiative for Hiring Ordinance, and San Francisco Fair Chance Ordinance, qualified applicants will be considered for assignment with arrest and conviction records. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness, meet client expectations, standards, and accompanying requirements, and safeguard business operations and company reputation. #TMMT