We are hiring a strong, hands-on data engineer who works across the full stack, from source ingestion and pipelines through modeling, governance, AI integration, and the reporting layer. You will design, build, and operate the platform yourself. We may add engineers later; the standards you set are the ones they will inherit.
Judgment is the job. You weren't hired just to execute. Be judicious. Weigh the cost-benefit against risk on every build-versus-buy decision and every shortcut.
Ship with discipline. Every pipeline, model, and metric ships with tests, documentation, version control, and a named owner.
Influence over authority. Much of the work depends on partners across the business. You earn their partnership. Run your big calls past me, then move.
What you'll do
Ingestion. Design, build, and operate batch and streaming pipelines from SaaS applications, databases, APIs, and custom in-house systems, using managed connectors where they fit and custom code (including change data capture) where they do not.
Lakehouse and governance. Build and administer a Databricks lakehouse (medallion architecture, Unity Catalog), including row filters, column masks, tagging, lineage, and identity-provider group sync.
Modeling and semantic layer. Develop dimensional models, entity resolution, and a governed metrics layer in dbt. Own testing, documentation, and data quality.
Unstructured data and ML. Build pipelines that turn audio, video, documents, images, and sensor or edge-device telemetry into structured data using transcription, OCR, LLM classification, and computer vision. Deploy, version, and monitor models with MLflow; build and maintain vector search indexes.
AI integration. Build and maintain LLM integrations (MCP servers, retrieval, text-to-SQL) with both read and write capabilities, enforcing user-level permissions, approval workflows, and audit logging.
APIs. Define integration requirements for APIs built by our application engineering team, and build against them.
Reporting. Build governed dashboards, reports, scheduled delivery, and alerting on a semantic-model BI platform.
Observability. Monitor pipeline freshness, data quality, model performance, and index health, with alerting that surfaces problems early.
Qualifications
8+ years in data engineering, including at least 3 owning a production cloud data platform end to end, from ingestion through the reporting layer.
Hands-on production experience with Databricks: Unity Catalog, Delta, Workflows/Lakeflow Jobs, and SQL warehouses.
Expert dbt and dimensional modeling: layered projects, tests, documentation, and a semantic or metrics layer (MetricFlow, Metric Views, or LookML).
Strong SQL and Python. Comfortable with Git, CI/CD, and infrastructure as code.
Built governed BI on a semantic model (Omni, Sigma, Looker, or similar) with row-level security, scheduled delivery, and alerting.
Delivered both managed-connector ingestion (Lakeflow Connect, Fivetran, Airbyte) and custom pipelines, including CDC and API sources.
Shipped at least one pipeline that turns unstructured data (audio, documents, or images) into structured data using transcription, OCR, LLM classification, or vision models.
Operated models in production with MLflow or an equivalent registry, including versioning, scoring jobs, and monitoring.
Implemented data access controls: row filters, column masks, PII handling, and group-based permissions synced from an identity provider.
Solved entity resolution across multiple source systems (customers, locations, or employees).
Built integrations that write back to production systems (in-house applications, CMS, or SaaS APIs) with authorization checks, idempotency, rollback, and audit logging.
Consumed APIs built by other teams and written clear requirements for what a data or AI integration needs from them.
Able to work directly with senior business leaders to define metrics and explain trade-offs in plain language.
Must be located around Glendale, CA as this role is 5 days on-site
Strongly Preferred
Built LLM applications on governed data: MCP servers, text-to-SQL with an evaluation set, and retrieval systems covering embedding models, chunking strategy, hybrid (keyword plus vector) search, metadata and permission filtering, and retrieval evaluation, using Databricks Vector Search, pgvector, or similar.
Experience with enterprise search or knowledge platforms (Glean or similar), knowledge graphs, or business glossaries.
Data observability tooling (Metaplane, Elementary, Monte Carlo, Lakehouse Monitoring).
Change-management automation: preview environments, automated tests, approval workflows, and policy-as-code (OPA, Cedar, or similar).
Streaming and IoT data: sensor or edge-device telemetry landed through Structured Streaming, Kafka, or edge-to-cloud pipelines.
Multi-cloud work across AWS, Azure, and GCP.
Marketing and advertising data (Google Ads, Meta, GA4, call tracking).
Home services, franchise, field service, or retail data.
Working knowledge of SOC 2 controls and California privacy requirements (CCPA/CPRA).
Benefits / Perks
We believe in recognizing and rewarding our employees for a job well done. We offer growth potential for motivated individuals, competitive compensation, and a comprehensive benefits package, including:
Medical, Dental, Vision, Life Insurance
401K Retirement Plan
Paid Vacation Time
Paid Holidays
and More!
Numbers & Facts
Location
Glendale, Arizona
Skills
Access Controlunmatched
Advertisingunmatched
Amazon Web Services (AWS)unmatched
Application Programming Interface (API)unmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Business Intelligenceunmatched
Call Monitoringunmatched
Centers for Disease Control and Prevention (CDC)unmatched
Cisco Unityunmatched
Cloud Computingunmatched
Content Management Systems (CMS)unmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Data Collectionunmatched
Data Qualityunmatched
Dimensional Modelingunmatched
Document Imagingunmatched
Documentationunmatched
GCP (Good Clinical Practices)unmatched
Gitunmatched
Internet of Thingsunmatched
Lookerunmatched
MCP - Microsoft Certified Professionalunmatched
Machine Toolunmatched
Marketingunmatched
Metadataunmatched
Metricsunmatched
Microsoft Windows Azureunmatched
On Site Supportunmatched
Performance Modelingunmatched
Privacy Controlsunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Reporting Dashboardsunmatched
Requirements Managementunmatched
Retailunmatched
Riskunmatched
SQL (Structured Query Language)unmatched
Shallow Parsingunmatched
Software Engineeringunmatched
Software as a Service (SaaS)unmatched
Structured Dataunmatched
Systems Administration/Managementunmatched
Team Lead/Managerunmatched
Telemetryunmatched
Test Automationunmatched
Unstructured Dataunmatched
Warehousingunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.