JD : · Establish prompt versioning, testing, and governance practices. · Improve consistency of agent behavior across enterprise use cases. Retrieval-Augmented Generation · Design and implement enterprise RAG architectures. · Build retrieval pipelines using enterprise documents, knowledge repositories, structured data, and metadata. · Optimize chunking, embedding, indexing, ranking, reranking, and retrieval strategies. · Improve grounding, citation quality, precision, recall, and factual accuracy. · Build reusable retrieval services for multiple agents and business domains. · Partner with data, platform, and knowledge management teams to onboard trusted enterprise knowledge sources. LLM and SLM Model Engineering · Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs for enterprise use cases. · Support domain-specific model development using internal and approved datasets. · Build supervised fine-tuning and model adaptation pipelines. · Apply model optimization techniques such as LoRA, QLoRA, distillation, quantization, and model compression. · Evaluate commercial, open-source, and internally hosted models for suitability, quality, cost, and operational fit. · Support model selection strategies based on use case sensitivity, latency, accuracy, cost, and data residency requirements. Private AI and On-Prem Model Hosting · Build and support private AI capabilities for hosting SLMs and LLMs in enterprise-controlled environments. · Deploy models on on-prem, hybrid, and private cloud infrastructure. · Support GPU-enabled model hosting using enterprise AI infrastructure. · Optimize model serving for latency, throughput, concurrency, resiliency, and GPU utilization. · Build secure inference endpoints for internal agent and application consumption. · Support air-gapped or restricted AI environments where required by security or compliance needs. · Partner with infrastructure and platform teams to operationalize private model hosting patterns. Model Serving and Inference Optimization · Implement scalable model serving using modern inference frameworks. · Build high-availability inference patterns for production workloads. · Optimize inference performance, token throughput, response latency, and cost efficiency. · Implement model routing, load balancing, caching, and fallback strategies. · Support batch inference and real-time inference use cases. · Develop reusable deployment templates for multiple model families and serving patterns. LLMOps, ModelOps, and AgentOps
· Build operational practices for managing models and agents across the lifecycle. · Implement observability for prompts, retrieval, model responses, latency, cost, and failures. · Develop evaluation pipelines for regression testing and continuous quality improvement. · Monitor model drift, response quality, hallucination indicators, and safety risks. · Support CI/CD and release management for prompts, models, agents, and retrieval pipelines. · Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness. AI Evaluation and Benchmarking · Define and implement LLM evaluation frameworks. · Measure accuracy, groundedness, relevance, hallucination rate, toxicity risk, safety compliance, task completion, and user satisfaction. · Build automated test suites for prompts, agents, tools, and RAG pipelines. · Benchmark models across enterprise use cases. · Compare cloud-hosted, open-source, and on-prem models based on performance, cost, quality, and risk. · Support go/no-go quality gates for production AI releases. Responsible AI, Security, and Governance · Embed Responsible AI controls into LLM applications and agent workflows. · Implement guardrails for safe output, tool usage, data access, and enterprise policy compliance. · Support model risk management, auditability, transparency, and traceability. · Ensure sensitive data is handled according to enterprise security and privacy requirements. · Partner with Security, Enterprise Architecture, Risk, and Compliance teams. · Support governance workflows for model approval, agent approval, and production readiness.