Role: Senior Engineer
Required Skills: Artificial Intelligence
Qualification: BACHELOR OF COMPUTER SCIENCE
Experience: 10 - 18 Years
Function: TECHNOLOGY
Apply By: 2026-11-30
Must Have Technical/Functional Skills
The Generative AI Architect will lead the architecture, design and hands-on technical delivery of secure, scalable, production-ready Generative AI solutions on AWS or Azure . This is a hands-on role. The architect is expected to write production-grade Python, build agentic workflows in LangGraph and LangChain, instrument solutions with Langfuse, and personally build proofs of concept, reference implementations and core framework components.
The architect will turn business requirements into reusable solution patterns. These cover large language models, Retrieval-Augmented Generation (RAG), Agentic AI, prompt orchestration, data pipelines, integrations, Responsible AI, observability and cloud-native deployment. The role requires the ability to evaluate open and closed foundation models. It also requires working with business stakeholders, enterprise architects, security teams, data teams, AI engineers, application teams and cloud platform teams throughout the solution lifecycle.
Roughly 40-50% of this role is hands-on engineering: building agent graphs, RAG pipelines, MCP tools, evaluation harnesses and deployment pipelines. The candidate must be able to prove this in a live coding or architecture-build exercise.
- 3+ years of hands-on experience designing and delivering AI, machine learning or Generative AI solutions.
- Proven experience taking Generative AI solutions from proof of concept to production deployment.
- Expert-level hands-on Python (async programming, typing, Pydantic, packaging, testing with pytest) and a recent, verifiable record of writing production code, not only designing solutions.
- Hands-on experience with Python agentic AI frameworks:
o LangGraph: stateful multi-agent graphs, checkpointers, interrupts and human-in-the-loop, memory stores.
o LangChain: LCEL, retrievers, tools, output parsers, integrations.
o Langfuse: tracing, prompt management, evaluations, datasets, cost analytics.
- Working knowledge of one or more additional frameworks: LlamaIndex, CrewAI, AutoGen/AG2, Semantic Kernel, OpenAI Agents SDK, Google ADK, Strands Agents or PydanticAI.
- Hands-on experience with Model Context Protocol (MCP), both building MCP servers and consuming MCP tools in agents, and familiarity with agent-to-agent (A2A) protocols.
- Strong architecture experience with at least one cloud provider (AWS, Microsoft Azure or Google Cloud Platform), and working knowledge of the comparable Generative AI, data, integration, security, monitoring and deployment services on the others. Relevant services include Amazon Bedrock, Azure OpenAI and Azure AI Foundry, and Vertex AI.
- Strong understanding of large language models, foundation models, transformers, tokenization, embeddings, prompt engineering, context windows, inference parameters, model limitations and evaluation techniques.
- Hands-on experience with RAG, vector databases, semantic and hybrid search, document chunking, metadata enrichment, reranking, grounding and citations. Relevant stores include pgvector, OpenSearch, Azure AI Search, Pinecone, Weaviate, Qdrant and Milvus.
- Experience with GraphRAG and knowledge graphs, for example Neo4j, is preferred.
- Experience designing and building agentic AI solutions with tool calling, orchestration, memory, planning, workflow control, approval gates and enterprise integrations.
- Experience designing REST APIs (FastAPI preferred), microservices, asynchronous integrations, event-driven systems and secure enterprise interfaces.
- Experience with relational databases, NoSQL databases, object storage, data lakes, search platforms and vector stores.
- Hands-on knowledge of container (Docker, Kubernetes), serverless and managed-platform deployment models.
- Understanding of Infrastructure as Code (Terraform, CloudFormation or Bicep), CI/CD, automated testing including LLM evaluation gates, release management and environment promotion.
- Strong understanding of cloud security, zero-trust principles, identity and access management, private networking, encryption, secrets management, audit logging and data governance.
- Experience defining non-functional requirements for availability, resiliency, scalability, performance, accessibility, maintainability, observability and cost.
- Excellent stakeholder management, consulting, facilitation, communication, documentation and presentation skills.
- Ability to explain complex Generative AI concepts, trade-offs, risks and architecture decisions to both technical and executive stakeholders.
Roles & Responsibilities
- Lead discovery workshops to understand business problems, user journeys, data availability, constraints, expected outcomes and measurable success criteria.
- Assess and prioritize Generative AI use cases by business value, technical feasibility, data readiness, security risk, implementation complexity and operating cost.
- Define high-level and low-level architectures for Generative AI applications on AWS, Azure or GCP.
- Design enterprise patterns for conversational AI, enterprise search, document intelligence, content generation, summarization, copilots, workflow automation and Agentic AI.
- Select foundation models, embedding models, orchestration frameworks, vector stores, data services and application components.
- Build agentic workflows hands-on in Python using LangGraph and LangChain. This includes stateful graphs, conditional routing, checkpointing and persistence, streaming, sub-graphs, tool binding, structured outputs and human-in-the-loop interrupts.
- Develop and maintain reusable Python agent frameworks, SDKs and accelerators. These include agent templates, tool and connector libraries, MCP servers and clients, prompt templates and evaluation harnesses for reuse across delivery teams.
- Implement LLM observability and evaluation with Langfuse, or an equivalent such as LangSmith, Arize Phoenix or OpenTelemetry-based tracing. This covers traces, spans, prompt versioning, cost and token tracking, user feedback capture, datasets, and LLM-as-judge and automated evaluations.
- Design RAG solutions covering ingestion, parsing, chunking, metadata enrichment, embeddings, indexing, hybrid search, reranking, grounding, source citation and response generation. Build them hands-on using LangChain or LlamaIndex.
- Decide when to use prompt engineering, RAG, GraphRAG, Retrieval-Augmented Fine-Tuning (RAFT), Parameter-Efficient Fine-Tuning (PEFT) or full model fine-tuning.
- Architect and build single-agent and multi-agent solutions. These should include tool use, short- and long-term memory, planning, supervisor/worker and hierarchical agent patterns, Model Context Protocol (MCP) integrations, approval gates and human-in-the-loop controls.
- Establish reusable prompt libraries, model-routing strategies, orchestration patterns, evaluation datasets and configuration-management practices.
- Design cloud-native deployment architectures with the right controls for scalability, availability, resiliency, disaster recovery, performance and cost.
- Build and expose agent services with FastAPI and async Python, containerize them with Docker and deploy them to Kubernetes, serverless or managed agent platforms such as Amazon Bedrock AgentCore, Azure AI Foundry Agent Service and Vertex AI Agent Engine.
- Build Responsible AI controls for content safety, bias, privacy, transparency, explainability, prompt injection, jailbreaks, hallucination, data leakage and inappropriate model behavior, using guardrail frameworks such as NeMo Guardrails, Guardrails AI or the native guardrails from each cloud provider.
- Define evaluation frameworks and quality metrics for relevance, groundedness, factuality, retrieval quality, hallucination rate, toxicity, latency, throughput, adoption and cost. Implement them with tools such as RAGAS, DeepEval and Langfuse evaluations.
- Establish LLMOps and GenAIOps practices across the lifecycle of models, prompts, agents, data, embeddings, vector indexes and applications. This includes CI/CD-integrated evaluation gates.
- Define observability for application health, agent execution, model performance, retrieval quality, token usage, latency, failures, security events and cloud consumption.
- Produce architecture documents, decision records, reference architectures, sequence diagrams, data-flow diagrams, deployment views, security models and operational runbooks.
- Lead hands-on through proof of concept, minimum viable product (MVP), pilot, production rollout and post-production stabilization. This includes writing code, reviewing pull requests and debugging production issues.
- Work with Product Owners and Scrum Masters to refine epics, features, user stories, acceptance criteria, dependencies and delivery plans.
- Provide technical leadership, mentoring, code and design reviews, design governance, troubleshooting and technical risk management across the Generative AI delivery pod.
- Ensure solutions comply with enterprise architect