Sr. Data Engineer | USC Only | W2 Contract | REMOTE.

Xlysi

  • Chicago, Illinois
  • 3 days ago
  • Remote
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Apache Sparkunmatched
    • Application Programming Interface (API)unmatched
    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Data Formatsunmatched
    • Data Lakeunmatched
    • Data Managementunmatched
    • Data Partitioningunmatched
    • Data Processingunmatched
    • Data Qualityunmatched
    • Data Scienceunmatched
    • Data Storageunmatched
    • Data Warehousingunmatched
    • Database Designunmatched
    • Database Extract Transform and Load (ETL)unmatched
    • Database Technologyunmatched
    • Distributed Computingunmatched
    • Distributed Databasesunmatched
    • Electronic Medical Recordsunmatched
    • Gitunmatched
    • Linux Operating Systemunmatched
    • Natural Language Parsingunmatched
    • NoSQLunmatched
    • Python Programming/Scripting Languageunmatched
    • REST (Representational State Transfer)unmatched
    • Relational Databases (RDBMS)unmatched
    • SNMP (Simple Network Management Protocol)unmatched
    • SQL Databasesunmatched
    • Scala Programming Languageunmatched
    • Scalable System Developmentunmatched
    • Source Code/Configuration Management (SCM)unmatched
    • Team Playerunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Test Dataunmatched
    • Unix Shell Programmingunmatched
    • Validation Testingunmatched

    Description

    Responsibilities:

    • Design and develop scalable ETL pipelines using Apache Spark (Scala/Python)
    • Build ingestion pipelines for network data (telemetry, syslogs, SNMP, configs, ticketing systems)
    • Transform raw data into query-ready formats in Data Lake
    • Implement monitoring, alerting, and pipeline reliability solutions
    • Develop CI/CD pipelines for data engineering workflows
    • Optimize large-scale data processing (batch & mini-batch, billions of events/day)
    • Manage data storage across distributed systems, databases, and APIs
    • Implement data quality checks, validation rules, and automated testing
    • Design and manage schemas, partitioning, and storage formats (Parquet, etc.)
    • Support data backfills and reprocessing for upstream or schema changes
    • Collaborate with data science and engineering teams for AI/ML data needs
    • Document data processes and ensure best practices across teams

    Requirements:

    • Strong experience with Apache Spark (Scala or Python)
    • Hands-on experience building ETL pipelines at scale
    • Strong Python skills (Spark/PySpark preferred)
    • Experience with AWS (S3, Glue, Athena, EMR)
    • Strong SQL and relational database knowledge
    • Experience with Airflow or similar orchestration tools
    • Knowledge of data warehousing, partitioning, and columnar storage (Parquet)
    • Experience with data quality frameworks and validation
    • Linux and shell scripting experience
    • Experience with Git and version control
    • Strong communication and collaboration skills

    Preferred:

    • Experience with Spark Streaming / Structured Streaming
    • Experience with Kafka or similar streaming platforms
    • NoSQL database experience
    • Experience with network data (syslogs, SNMP, telemetry)
    • REST API / cloud SDK integration (boto3, etc.)
    • Experience with automated testing for data pipelines
    • Knowledge of log parsing / text analytics
    • Telecom or large-scale network environment experienceS

    Numbers & Facts

    LocationChicago, Illinois (
    Remote
    )
    Websitehttp://www.xlysi.com

    Similar Jobs