Operate always-on model inference capacity , including monitoring traffic, resolving throttling issues, tuning auto-scaling, requesting additional capacity, redeploying services, and escalating issues to platform owners as needed. This role will be responsible for the end-to-end lifecycle of ML models, including model evaluation, inference services, deployment, monitoring, troubleshooting, and integration into internal tools and workflows.