• Glendale, Arizona
    4 days ago

    Job Description

    About the role

    We are hiring a strong, hands-on data engineer who works across the full stack, from source ingestion and pipelines through modeling, governance, AI integration, and the reporting layer. You will design, build, and operate the platform yourself. We may add engineers later; the standards you set are the ones they will inherit.

    • Judgment is the job. You weren't hired just to execute. Be judicious. Weigh the cost-benefit against risk on every build-versus-buy decision and every shortcut.
    • Ship with discipline. Every pipeline, model, and metric ships with tests, documentation, version control, and a named owner.
    • Influence over authority. Much of the work depends on partners across the business. You earn their partnership. Run your big calls past me, then move.


    What you'll do

    • Ingestion. Design, build, and operate batch and streaming pipelines from SaaS applications, databases, APIs, and custom in-house systems, using managed connectors where they fit and custom code (including change data capture) where they do not.
    • Lakehouse and governance. Build and administer a Databricks lakehouse (medallion architecture, Unity Catalog), including row filters, column masks, tagging, lineage, and identity-provider group sync.
    • Modeling and semantic layer. Develop dimensional models, entity resolution, and a governed metrics layer in dbt. Own testing, documentation, and data quality.
    • Unstructured data and ML. Build pipelines that turn audio, video, documents, images, and sensor or edge-device telemetry into structured data using transcription, OCR, LLM classification, and computer vision. Deploy, version, and monitor models with MLflow; build and maintain vector search indexes.
    • AI integration. Build and maintain LLM integrations (MCP servers, retrieval, text-to-SQL) with both read and write capabilities, enforcing user-level permissions, approval workflows, and audit logging.
    • APIs. Define integration requirements for APIs built by our application engineering team, and build against them.
    • Reporting. Build governed dashboards, reports, scheduled delivery, and alerting on a semantic-model BI platform.
    • Observability. Monitor pipeline freshness, data quality, model performance, and index health, with alerting that surfaces problems early.


    Qualifications

    • 8+ years in data engineering, including at least 3 owning a production cloud data platform end to end, from ingestion through the reporting layer.
    • Hands-on production experience with Databricks: Unity Catalog, Delta, Workflows/Lakeflow Jobs, and SQL warehouses.
    • Expert dbt and dimensional modeling: layered projects, tests, documentation, and a semantic or metrics layer (MetricFlow, Metric Views, or LookML).
    • Strong SQL and Python. Comfortable with Git, CI/CD, and infrastructure as code.
    • Built governed BI on a semantic model (Omni, Sigma, Looker, or similar) with row-level security, scheduled delivery, and alerting.
    • Delivered both managed-connector ingestion (Lakeflow Connect, Fivetran, Airbyte) and custom pipelines, including CDC and API sources.
    • Shipped at least one pipeline that turns unstructured data (audio, documents, or images) into structured data using transcription, OCR, LLM classification, or vision models.
    • Operated models in production with MLflow or an equivalent registry, including versioning, scoring jobs, and monitoring.
    • Implemented data access controls: row filters, column masks, PII handling, and group-based permissions synced from an identity provider.
    • Solved entity resolution across multiple source systems (customers, locations, or employees).
    • Built integrations that write back to production systems (in-house applications, CMS, or SaaS APIs) with authorization checks, idempotency, rollback, and audit logging.
    • Consumed APIs built by other teams and written clear requirements for what a data or AI integration needs from them.
    • Able to work directly with senior business leaders to define metrics and explain trade-offs in plain language.
    • Must be located around Glendale, CA as this role is 5 days on-site


    Strongly Preferred

    • Built LLM applications on governed data: MCP servers, text-to-SQL with an evaluation set, and retrieval systems covering embedding models, chunking strategy, hybrid (keyword plus vector) search, metadata and permission filtering, and retrieval evaluation, using Databricks Vector Search, pgvector, or similar.
    • Experience with enterprise search or knowledge platforms (Glean or similar), knowledge graphs, or business glossaries.
    • Data observability tooling (Metaplane, Elementary, Monte Carlo, Lakehouse Monitoring).
    • Change-management automation: preview environments, automated tests, approval workflows, and policy-as-code (OPA, Cedar, or similar).
    • Streaming and IoT data: sensor or edge-device telemetry landed through Structured Streaming, Kafka, or edge-to-cloud pipelines.
    • Multi-cloud work across AWS, Azure, and GCP.
    • Marketing and advertising data (Google Ads, Meta, GA4, call tracking).
    • Home services, franchise, field service, or retail data.
    • Working knowledge of SOC 2 controls and California privacy requirements (CCPA/CPRA).


    Benefits / Perks

    We believe in recognizing and rewarding our employees for a job well done. We offer growth potential for motivated individuals, competitive compensation, and a comprehensive benefits package, including:

    • Medical, Dental, Vision, Life Insurance
    • 401K Retirement Plan
    • Paid Vacation Time
    • Paid Holidays
    • and More!

    Numbers & Facts

    LocationGlendale, Arizona

    Skills

    • Access Controlunmatched
    • Advertisingunmatched
    • Amazon Web Services (AWS)unmatched
    • Application Programming Interface (API)unmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Business Intelligenceunmatched
    • Call Monitoringunmatched
    • Centers for Disease Control and Prevention (CDC)unmatched
    • Cisco Unityunmatched
    • Cloud Computingunmatched
    • Content Management Systems (CMS)unmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Data Collectionunmatched
    • Data Qualityunmatched
    • Dimensional Modelingunmatched
    • Document Imagingunmatched
    • Documentationunmatched
    • GCP (Good Clinical Practices)unmatched
    • Gitunmatched
    • Internet of Thingsunmatched
    • Lookerunmatched
    • MCP - Microsoft Certified Professionalunmatched
    • Machine Toolunmatched
    • Marketingunmatched
    • Metadataunmatched
    • Metricsunmatched
    • Microsoft Windows Azureunmatched
    • On Site Supportunmatched
    • Performance Modelingunmatched
    • Privacy Controlsunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reporting Dashboardsunmatched
    • Requirements Managementunmatched
    • Retailunmatched
    • Riskunmatched
    • SQL (Structured Query Language)unmatched
    • Shallow Parsingunmatched
    • Software Engineeringunmatched
    • Software as a Service (SaaS)unmatched
    • Structured Dataunmatched
    • Systems Administration/Managementunmatched
    • Team Lead/Managerunmatched
    • Telemetryunmatched
    • Test Automationunmatched
    • Unstructured Dataunmatched
    • Warehousingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder