Join a Platform Engineering team responsible for operating and supporting distributed caching services used by critical business applications. This role focuses on the day-to-day administration, monitoring, troubleshooting, and maintenance of Redis-based caching platforms running in enterprise environments. Working alongside senior engineers, you will help ensure platform reliability, performance, and availability while contributing to automation and operational improvements. Experience with container platforms is beneficial but not required.
Key Responsibilities
- Support daily operations of Redis and caching platforms in production and non-production environments
- Monitor platform health, performance, capacity, and availability
- Investigate and resolve operational incidents, alerts, and service disruptions
- Execute maintenance activities including upgrades, patching, backups, and recovery procedures
- Assist with Redis configuration, deployment, and troubleshooting activities
- Develop and maintain operational automation using Ansible and scripting tools
- Create and maintain operational documentation, runbooks, and knowledge articles
- Collaborate with application teams to support caching-related performance and connectivity issues
- Participate in on-call and operational support rotations as required
- Support platform modernization initiatives, including containerized deployments when applicable
Basic Qualifications
- Experience supporting Linux-based systems in enterprise environments
- Understanding of system administration fundamentals including networking, processes, storage, and troubleshooting
- Experience supporting production applications or infrastructure platforms
- Exposure to Redis, caching technologies, or similar distributed data platforms
- Experience with Ansible or similar automation and configuration management tools
- Familiarity with Bash, Python, or similar scripting languages
- Strong analytical and problem-solving skills
- Ability to document procedures and work effectively within operational support processes
- Good communication and teamwork skills
Preferred Qualifications
- Experience with Redis administration or support
- Understanding of caching concepts such as key-value stores, data expiration, replication, and performance tuning
- Exposure to Kubernetes, containers, Docker, or Helm
- Experience with monitoring and observability tools
- Experience supporting high-availability systems
- Relevant certifications, coursework, or equivalent practical experience
Accountabilities
- Designs, codes, tests, debugs, and documents software according to systems quality standards, policies, and procedures
- Analyzes business needs and creates software solutions
- Responsible for preparing design documentation
- Prepares test data for unit, string, and parallel testing
- Evaluates and recommends software and hardware solutions to meet user needs
- Resolves customer issues with software solutions and responds to suggestions for improvements and enhancements
- Works with business and development teams to clarify requirements to ensure testability
- Drafts, revises, and maintains test plans, test cases, and automated test scripts
- Executes test procedures according to software requirements specifications
- Logs defects and makes recommendations to address defects
- Retests software corrections to ensure problems are resolved
- Documents evolution of testing procedures for future replication
- May conduct performance and scalability testing
Responsibilities
- Plans, conducts, and leads assignments generally involving moderate, high-budget projects or more than one project
- Manages user expectations regarding appropriate milestones and deadlines
- Assists in training, work assignment, and checking of less experienced developers
- Serves as technical consultant to leaders in the IT organization and functional user groups
- Subject matter expert in one or more technical programming specialties; employs expertise as a generalist or a specialist
- Performs estimation efforts on complex projects and tracks progress
- Works on the highest level of problems where analysis of situations or data requires an in-depth evaluation of various factors
- Documents, evaluates, and researches test results; documents evolution of testing scripts for future replication
- Identifies, recommends, and implements changes to enhance the effectiveness of quality assurance strategies
Primary responsibility will be for the day-to-day operation, maintenance, and reliability of enterprise caching services, including Redis and related distributed caching technologies, deployed across diverse infrastructure environments. Key duties will include monitoring platform health and performance, executing planned maintenance activities such as patching, upgrades, and capacity adjustments, troubleshooting incidents, and driving root cause analysis.
You will contribute to the development and enhancement of automation frameworks to streamline operational workflows, reduce manual intervention, and improve service consistency. Collaboration with engineering, application, and infrastructure teams will be essential to ensure caching services meet performance, availability, and scalability requirements across the enterprise.