Job Title: AI Technical Architect
Location: Dallas, TX, - Hybrid
Key Responsibilities
Define AI/ML reference architecture and solution blueprints (batch/streaming ML, LLM + RAG, multimodal).
Lead end-to-end solution design across data ingestion, model training, inference, deployment, and monitoring.
Architect LLM applications (agents, summarization, classification) with RAG, evaluation frameworks, safety controls, and guardrails.
Own MLOps/LLMOps practices, including CI/CD for models, model registry, feature stores, lineage tracking, observability, drift detection, and cost monitoring.
Choose the right cloud and runtime strategy (managed services vs. self-hosted, GPU vs. CPU, serverless vs. containerized).
Establish AI governance standards, including PII handling, encryption, auditability, and Responsible AI practices.
Collaborate with product and business stakeholders to translate requirements into architectural decisions and delivery plans.
Perform technical spikes and POCs, benchmark models and infrastructure, and lead Architecture Reviews.
Create and maintain standards, patterns, and reusable components; mentor engineers across teams.
Drive performance and cost optimization initiatives, including throughput, latency, SLA/SLO management, caching, quantization/distillation, and autoscaling.
Support vendor and product evaluations, including cloud AI services, vector databases, orchestration frameworks, and monitoring platforms.
Required Qualifications
Bachelor’s or master’s degree in computer science, Engineering, Data Science, AI, or a related field.
15+ years of overall engineering experience, with at least 4+ years in AI/ML solution architecture.
Proven experience designing and deploying AI systems in production at scale (LLM and/or classical ML).
Strong hands-on proficiency in Python and at least one major cloud platform (AWS, Azure, or GCP).
Must-Have Technical Skills
AI/ML
LLM Architecture
Designing LLM/RAG systems, including retrieval pipelines, chunking strategies, embeddings, reranking, prompt orchestration, response orchestration, evaluation, and safety.
Deep understanding of the model lifecycle, including fine-tuning, PEFT/LoRA, quantization, distillation, latency optimization, and cost optimization.
Strong ML/NLP expertise, including feature engineering, model selection, training, cross-validation, experimentation, and testing.
MLOps / LLMOps
CI/CD for ML, including model versioning, model promotion, feature stores, model registry, lineage tracking, and drift detection.
Inference stacks including PyTorch, TensorFlow, vLLM, TGI, ONNX, GPU orchestration, autoscaling, and APM.
Pipelines and orchestration frameworks such as Airflow, Kubeflow, and MLflow.