Lead Data Architect

Karsun Solutions, LLC
  • Herndon, Virginia
  • Full-time
30+ days ago

Job Description

Overview:

Summary

The Lead Data Architect will design, build, and operate enterprise data platforms that power GenAI and AI/ML use cases. This is a highly technical, hands-on role responsible for data platform architecture, end-to-end data engineering, ML/LLM pipeline design, production model onboarding, and delivery of scalable Databricks- centric solutions across cloud environments.

Responsibilities:

What You'll Be Doing:

  • Architect and implement enterprise data platforms (batch + streaming) optimized for ML, LLMs, and GenAI workloads.
  • Lead design and hands on implementation of Databricks workspaces, Unity Catalog, Delta Lake design patterns, cluster policies, and performance tuning.
  • Build and own end to end data pipelines (ingest, transform, feature engineering, serving) using PySpark, Databricks Jobs, Spark SQL, Delta Lake, and orchestration tools.
  • Design and operationalize model training, fine tuning (LLM), evaluation, deployment, and monitoring pipelines (MLOps/RAG/CAG) integrating Databricks MLflow, CI/CD, and infra-as-code.
  • Implement vectorless and vectorization/embedding pipelines, vector store integrations, and retrieval layers for RAG (FAISS, Pinecone, Weaviate, Milvus).
  • Define data schemas, governance, lineage, access controls, and data product APIs; implement Unity Catalog or equivalent for centralized governance.
  • Drive cost/performance optimization for storage, compute (spot/preemptible),and query patterns. 
  • Collaborate with engineers, data scientists, product owners, and security to translate business needs into production GenAI solutions. 
  • Mentor and lead engineering teams; conduct architecture reviews, code reviews, and run technical deep dives. 
  • Implement observability for data and ML pipelines (metrics, logging, data quality tests, alerting). 
  • Create reproducible experiment tracking, model registry, and rollout strategies (canary, shadow testing, rollback). 
  • Stay current on GenAI/LLM architectures and evaluate/introduce new tooling and frameworks.
Qualifications and Education:

Required Qualifications:

  • BA or BS degree in CS, Computer Engineering, Information Technology or a
    related field.
  • 8+ years hands on experience in data engineering/platform architecture; 3+ years in an architect or lead role.
  • Candidate must hold an active AWS Certified Machine Learning – Specialty certification or equivalent AWS certification.
  • Proven, hands on Databricks experience (designing workspaces, Delta Lake, performance tuning, productionizing Spark jobs).
  • Deep Spark + PySpark expertise and experience with Databricks Runtime.
  • Strong experience building ML/LLM pipelines and operationalizing models (training, fine tuning, serving).
  • Practical experience with vector embeddings, semantic search, and RAG architectures.
  • Solid Python expertise and common ML libraries (PyTorch, TensorFlow, Hugging Face transformers) and MLflow.
  • Cloud platform experience (AWS strongly preferred).
  • Experience with containerization and orchestration while leveraging open source libraries for unstructured and structured data processing, serving/inference.
  • Strong SQL skills; experience with distributed query/warehouse systems and parquet/AVRO/Delta formats.
  • CI/CD and infra-as-code experience (Terraform, GitOps, Jenkins/GitHub Actions/GitLab CI).
  • Data governance, security, and IAM experience; experience implementing row/column level access controls and data lineage.
  • Demonstrated ability to design for scalability, reliability, and cost efficiency. 
Compensation:

The proposed salary range for this role is $****** to $******* USD. The salary range provided is a good faith estimate representative of all experience levels. Karsun considers several factors when extending an offer, including but not limited to, the role, function and associated responsibilities, a candidate’s work experience, location, education/training, and key skills.

Numbers & Facts

LocationHerndon, Virginia
Job TypeFull-time

Skills

  • Access Controlunmatched
  • Amazon Web Services (AWS)unmatched
  • Apache Avrounmatched
  • Application Programming Interface (API)unmatched
  • Architectural Analysisunmatched
  • Architectural Designunmatched
  • Artificial Intelligence (AI)unmatched
  • Cisco Unityunmatched
  • Cloud Computingunmatched
  • Code Reviewsunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Cost Controlunmatched
  • Data Managementunmatched
  • Data Processingunmatched
  • Data Qualityunmatched
  • Data Scienceunmatched
  • DataArchitect Data Modeling Toolunmatched
  • Design Patterns Programming Methodologiesunmatched
  • GitHubunmatched
  • Information/Data Security (InfoSec)unmatched
  • Jenkinsunmatched
  • Machine Learningunmatched
  • Machine Toolunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Onboardingunmatched
  • Open Sourceunmatched
  • Performance Tuning/Optimizationunmatched
  • Python Programming/Scripting Languageunmatched
  • SQL (Structured Query Language)unmatched
  • Semantic Searchunmatched
  • Structured Dataunmatched
  • Team Lead/Managerunmatched
  • Unstructured Dataunmatched
  • Use Casesunmatched
  • Warehousingunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder