Data / AI Native Engineer

Kasmo Inc
  • Arlington, NJ
  • Instant Apply
1 day ago

Job Description



Description:
Job Description Data / AI Native Engineer
Location: St. louis, MO/ Arlington, VA/ NYC, NY & NJ
Remote is OK | Arlington & NY/NJ is hybrid
Day-to-Day Job Duties:
Design, build, optimize, and support enterprise-scale Databricks and AWS Lakehouse data pipelines using Spark and PySpark.
Develop Databricks solutions leveraging Delta Lake, Medallion Architecture (Bronze/Silver/Gold), and Unity Catalog.
Develop high-volume batch and streaming pipelines using PySpark, ensuring scalability, performance, reliability, data quality, and maintainability.
Design and implement modern lakehouse engineering patterns including SCD Type 2, Change Data Capture (CDC), idempotent writes, schema evolution, and data contracts.
Work across the AWS data ecosystem, including technologies such as S3, Glue, Lambda, EMR, Redshift, Kinesis, and related AWS services.
Use GitHub Copilot and/or Claude Code CLI as part of the daily engineering workflow to accelerate application code, data-pipeline development, testing, debugging, optimization, and documentation.
Rapidly develop proofs of concept and MVPs using AI-assisted development, while establishing clear paths to scalable production implementation.
Apply AI capabilities directly within the data-engineering lifecycle, including use cases such as:
AI-assisted schema inference
Automated data-quality rule generation
LLM-based anomaly detection
AI-generated PySpark transformations
Intelligent data validation and enrichment
Build data foundations for AI-enabled products, including feature stores, vector stores, RAG data layers, embedding pipelines, and other AI-ready data architectures.
Design and develop high-volume streaming pipelines using technologies such as Spark Structured Streaming, Kafka, Kinesis, and Databricks Auto Loader.
Develop and orchestrate enterprise data pipelines using technologies such as Databricks Workflows, Delta Live Tables, Airflow, and dbt.
Implement secure and governed data-engineering practices using Unity Catalog, including catalog/schema organization, access control, lineage, governance, and data discovery.
Lead technically complex data-engineering initiatives, drive solution design, coordinate engineering activities, and help resolve complex platform and pipeline issues.
Partner with architects, software engineers, product owners, source-system teams, governance teams, analytics teams, AI teams, and business stakeholders.
Remain actively hands-on with PySpark, Databricks, AWS, and software-engineering activities, rather than operating solely as a people manager or high-level data architect.
Basic Qualifications:
Minimum 8+ years of overall Data Engineering / Software Engineering experience delivering enterprise-scale data solutions.
Minimum 5+ years of strong hands-on Data Engineering experience, including design, development, optimization, and support of production data pipelines.
Minimum 3+ years of hands-on Databricks experience, including production experience with:
Delta Lake
Medallion Architecture Bronze / Silver / Gold
Unity Catalog
Minimum 5+ years of hands-on Spark experience, with strong expertise in PySpark specifically, rather than Scala-only Spark development.
Minimum 3+ years of AWS data-platform experience, using services such as S3, Glue, Lambda, EMR, Redshift, Kinesis, or related AWS technologies.
Minimum 2+ years of experience leading complex data-engineering initiatives, technical workstreams, or engineering teams while remaining hands-on technically.
Proven experience combining data engineering and software-engineering practices, including modular development, testing, source control, CI/CD, observability, and production support.
Demonstrated daily hands-on use of GitHub Copilot and/or Claude Code CLI, with the ability to describe concrete examples of how these tools are used to accelerate engineering work.
Demonstrated ability to build POCs and MVPs rapidly using AI-assisted development.
Strong understanding of distributed data-processing principles, large-scale pipeline design, performance optimization, data quality, and production operations.
Strong SQL skills and experience working with large enterprise datasets.
Critical Screening Requirements:
Candidates should be screened out if they do not demonstrate:
Proven experience spanning Data Engineering and Software Engineering.
Proven experience leading complex data-engineering initiatives or teams.
Hands-on Databricks experience including Delta Lake, Bronze/Silver/Gold Medallion Architecture, and Unity Catalog.
High-volume Spark and PySpark experience; Scala-only Spark experience is not sufficient.
Strong experience with the AWS data ecosystem.
Daily hands-on use of GitHub Copilot and/or Claude Code CLI with specific examples.
Ability to use AI-assisted development to rapidly create POCs and MVPs.
Travel
Travel as required based on project and client needs.
Degree
Bachelor's degree in Computer Science, Data Engineering, Information Systems, Engineering, or a related technical discipline, or equivalent professional work experience.
Nice to Have (But Not a Must)
2+ years of experience applying AI/Generative AI within data-engineering solutions, particularly:
AI-assisted schema inference
Automated data-quality rules
LLM-based anomaly detection
AI-generated PySpark transformations
Experience building AI-enabled data products, including feature stores, vector stores, RAG data layers, and embedding pipelines.
2+ years of high-volume streaming experience using Spark Structured Streaming, Kafka, Kinesis, Auto Loader, or equivalent technologies.
Strong implementation experience with modern lakehouse patterns including:
SCD Type 2
CDC
Idempotent writes
Schema evolution
Data contracts
Hands-on experience with dbt, Airflow, Databricks Workflows, and/or Delta Live Tables.
Experience designing data platforms and pipelines supporting enterprise-scale GenAI, RAG, machine-learning, or agentic-AI solutions.
Profiles That Typically Align Well
Senior Data Engineer Lead Data Engineer Principal Data Engineer Data Platform Engineer Senior Analytics Engineer Data Engineering Manager with active individual-contributor responsibilities
Profiles That Typically Do Not Align
Traditional ETL Developer focused primarily on Informatica / Ab Initio / DataStage without cloud-native engineering BI / Report Developer Diagram-only Data Architect DBA Hadoop Administrator

Numbers & Facts

LocationArlington, NJ

Skills

  • AWS Lambdaunmatched
  • Access Controlunmatched
  • Amazon Simple Storage Service (S3)unmatched
  • Amazon Web Services (AWS)unmatched
  • Apache Hadoopunmatched
  • Artificial Intelligence (AI)unmatched
  • Business Intelligenceunmatched
  • Centers for Disease Control and Prevention (CDC)unmatched
  • Cisco Unityunmatched
  • Cloud Computingunmatched
  • Computer Scienceunmatched
  • Concreteunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Data Collectionunmatched
  • Data Managementunmatched
  • Data Qualityunmatched
  • Data Setsunmatched
  • DataArchitect Data Modeling Toolunmatched
  • Database Administrationunmatched
  • Database Extract Transform and Load (ETL)unmatched
  • Electronic Medical Recordsunmatched
  • Engineeringunmatched
  • Engineering Managementunmatched
  • GitHubunmatched
  • High Level Architecture (HLA)unmatched
  • IBM WebSphere DataStageunmatched
  • Informaticaunmatched
  • Information/Data Security (InfoSec)unmatched
  • Leadershipunmatched
  • Machine Learningunmatched
  • People Managementunmatched
  • Performance Tuning/Optimizationunmatched
  • Production Supportunmatched
  • Proof of Conceptunmatched
  • SQL (Structured Query Language)unmatched
  • Software Architectureunmatched
  • Software Engineeringunmatched
  • Source Code/Configuration Management (SCM)unmatched
  • Streaming Technologyunmatched
  • Systems Engineeringunmatched
  • Technical Leadershipunmatched
  • Test Plan/Scheduleunmatched
  • Use Casesunmatched
  • Willing to Travelunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder