Seeking a Data Scientist to support a Natural Language Processing (NLP) project focused on developing automated solutions to process, analyze, and annotate language data derived from both spoken and written sources
This role involves building and improving models that tokenize language data, assign linguistic features such as parts of speech, and evaluate model performance against human-generated annotations to ensure accuracy and reliability
Primary Responsibilities
Develop and implement NLP models and algorithms to tokenize and process spoken and written language data
Design and build automated annotation solutions for labeling language data, including parts-of-speech tagging and related linguistic features
Evaluate and improve model performance by comparing outputs against human-generated annotations and refining algorithms accordingly
Analyze large volumes of structured and unstructured language data to identify patterns and improve model accuracy
Support development of training datasets and annotation frameworks to enhance model performance
Collaborate with cross-functional teams to align NLP solutions with mission and data requirements
Perform data preprocessing, cleansing, and transformation to prepare datasets for modeling
Document methodologies, model performance metrics, and findings to support repeatability and continuous improvement
Stay current with advancements in NLP, machine learning, and language modeling techniques
Required Qualifications
Must have active Top Secret/SCI clearance with Full Scope Polygraph (MD Customer)
Master’s Degree with 6 years of relevant experience, Bachelor’s Degree with 8 years of relevant experience, or Associate's Degree with 10 years of in-depth relevant experience that is clearly related to the position
Experience with Natural Language Processing (NLP) techniques and tools
Proficiency in Python and common NLP/data science libraries (e.g., NLTK, spaCy, scikit-learn)
Experience with text and speech data processing, including tokenization and annotation
Ability to evaluate model performance using quantitative metrics and comparison methods
Strong analytical and problem-solving skills
Desired Qualifications
Experience working with annotated language datasets and linguistic frameworks
Familiarity with machine learning or deep learning models for NLP
Experience with speech processing or audio data analysis
Knowledge of model evaluation techniques and performance benchmarking
Experience working in mission-driven or government environments
Exempt hourly position. 11 paid holidays, minimum of 3 weeks PTO, company sponsored group medical plan, company paid dental, vision, life insurance, and STD/LTD plans. Salary is dependent upon the candidate’s experience and qualifications.
Numbers & Facts
Location
Maryland
Skills
Algorithmsunmatched
Analysis Skillsunmatched
Benchmarkingunmatched
Continuous Improvementunmatched
Cross-Functionalunmatched
Data Analysisunmatched
Data Modeling Languageunmatched
Data Processingunmatched
Data Scienceunmatched
Deep Learningunmatched
Documentation Modelsunmatched
Full Scope Polygraphunmatched
Governmentunmatched
Health Planunmatched
Life Insuranceunmatched
Linguisticsunmatched
Machine Learningunmatched
Modeling Languagesunmatched
Natural Language Processing (NLP)unmatched
Natural Language Toolkit (NLTK)unmatched
Performance Managementunmatched
Performance Metricsunmatched
Performance Modelingunmatched
Problem Solving Skillsunmatched
Python Programming/Scripting Languageunmatched
Science Libraryunmatched
Sensitive Compartmented Information (SCI)unmatched
Top Secret Clearanceunmatched
Training Data Setsunmatched
Unstructured Dataunmatched
Vision Planunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.