Prometheus, Grafana, DataDog) and logging frameworks for production systems Basic knowledge of machine learning concepts and MLOps, including data pipelines, model versioning, and experiment tracking toolsExpertise in open source data analytics and governance platforms: architecture, deployment, and performance tuning of Datahub, Apache Spark, Flink, Hive, Hadoop/HDFS, and Iceberg Rest Catalog Experience building multi-agent AI systems: proficiency with LangChain, LangGraph, or AutoGen frameworks; strong prompt engineering and LLM integration skills; ability to design event-driven architectures for autonomous workflows Skills in integration and communication layers: implement MCP servers and APIs using Python, REST/GraphQL, and message queuing (e.g. Kafka, RabbitMQ experience with modern data platforms including Snowflake, Databricks, and vector databases MLOps and observability capabilities: deploy containerized AI systems with comprehensive monitoring; track experiments using MLflow or Weights & Biases; implement distributed tracing for agent workflows and model performance.