Experience with workflow orchestration tools including Apache Airflow and Beam Experience with AWS: e.g., S3, EMR, Lambda, Glue, Redshift, BigQuery, Kinesis, or similar services Experience with Analytics frameworks including Trino (Presto, BigQuery, Snowflake) Hands-on experience with big data lake architectures Experience with containerization and orchestration (Docker, Kubernetes/EKS) and CI/CD tooling including Jenkins Experience in Python and PySpark Familiarity with graph databases such as TigerGraph Experience building pipelines that process multimodal data (structured and image) and integrate ML model inference - including LLMs and embedding models - for data enrichment and transformation Hands-on experience deploying, serving, and optimizing LLMs or ML models directly in the production, inference runtimes/compilers (ONNX Runtime, TensorRT/TensorRT-LLM), and serving frameworks (Triton, vLLM, TorchServe or similar). Experience tuning batching, KV-cache, and GPU utilization for low-latency, high-throughput real-time inference in a data pipeline Knowledge of data governance principles, data security best practices, and data privacy regulations Proven experience delivering a consumer-oriented solution by participating at every stage of the development life-cycle.