Seeking a Data Scientist to support the development of advanced Natural Language Processing (NLP) solutions focused on the automated processing and analysis of spoken and written language data
This role is responsible for building and improving models that accurately tokenize language data and annotate linguistic features, enabling scalable and high-quality language understanding capabilities
The position involves evaluating model performance against human-generated annotations to continuously enhance accuracy and effectiveness
Primary Responsibilities
Develop and implement NLP models to automatically tokenize language data from both spoken and written sources
Design and build automated solutions for annotating language data with parts-of-speech (POS) tagging and linguistic features
Train, test, and validate machine learning models for language processing tasks
Evaluate and improve model performance by benchmarking against human-generated annotations for speech and text data
Process and analyze large volumes of structured and unstructured language data
Develop data pipelines to support ingestion, preprocessing, and transformation of linguistic datasets
Collaborate with cross-functional teams to integrate NLP models into production systems
Perform error analysis and model tuning to improve accuracy and robustness
Document methodologies, model performance, and data processing workflows
Required Qualifications
Must have active Top Secret/SCI clearance with Full Scope Polygraph (MD Customer)
Master’s degree with 6 years of relevant experience, Bachelor’s Degree with 8 years of relevant experience, or Associate's Degree with 10 years of in-depth relevant experience that is clearly related to the position
Experience with Natural Language Processing (NLP) and machine learning techniques
Strong proficiency in Python and relevant NLP/ML libraries (e.g., NLTK, spaCy, Hugging Face, or similar)
Experience developing models for tokenization, text processing, and linguistic annotation
Experience working with speech and/or text data
Understanding of model evaluation techniques, including comparison to ground truth or human-labeled data
Strong analytical and problem-solving skills
Desired Qualifications
Experience with part-of-speech tagging and linguistic modeling
Experience working with multilingual or speech-based datasets
Familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow)
Experience deploying NLP models in production environments
Knowledge of data pipelines and large-scale data processing
Exempt hourly position. 11 paid holidays, minimum of 3 weeks PTO, company sponsored group medical plan, company paid dental, vision, life insurance, and STD/LTD plans. Salary is dependent upon the candidate’s experience and qualifications.
Numbers & Facts
Location
Maryland
Skills
Analysis Skillsunmatched
Benchmarkingunmatched
Cross-Functionalunmatched
Data Managementunmatched
Data Processingunmatched
Data Scienceunmatched
Data Setsunmatched
Deep Learningunmatched
Documentation Modelsunmatched
Full Scope Polygraphunmatched
Health Planunmatched
Life Insuranceunmatched
Linguisticsunmatched
Machine Learningunmatched
Model Validationunmatched
Modeling Languagesunmatched
Multilingualunmatched
Natural Language Processing (NLP)unmatched
Natural Language Toolkit (NLTK)unmatched
Performance Managementunmatched
Performance Modelingunmatched
Problem Solving Skillsunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Sensitive Compartmented Information (SCI)unmatched
Top Secret Clearanceunmatched
Training Data Setsunmatched
Unstructured Dataunmatched
Vision Planunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.