4 - 10 hour days (Sunday-Wednesday 7am-5pm)
Database Site Reliability Engineer (Database Operations)Position Summary
We are seeking an experienced Database Site Reliability Engineer (SRE) to support and operate mission-critical database platforms within a fast-paced enterprise environment. This role is focused on operational excellence, reliability, resiliency, automation, and continuous improvement across multiple database technologies.
This is not a development role. We are looking for hands-on database professionals who thrive in production operations, take ownership of issues, understand urgency, and are passionate about improving systems, processes, and themselves.
The ideal candidate views reliability as a product, proactively identifies risks before they become incidents, and continuously seeks opportunities to automate repetitive tasks and improve platform stability.
Key Responsibilities
Database Operations & Reliability
·Install, configure, upgrade, patch, and maintain enterprise database platforms.
Ensure availability, performance, recoverability, and security of production database environments.
Monitor database platforms and associated infrastructure, responding rapidly to incidents and service degradations.
Lead troubleshooting efforts for database, operating system, storage, replication, and application connectivity issues.
Execute failovers, disaster recovery testing, and recovery procedures.
Partner with application teams to provide database guidance and operational support.
Platform Engineering
·Deploy, maintain, and optimize database infrastructure across physical, virtual, and cloud environments.
Implement scalable, resilient database solutions.
Evaluate and recommend improvements to architecture, monitoring, automation, and operational processes.
Support capacity planning, performance tuning, and platform lifecycle management.
Automation & Continuous Improvement
·Develop and maintain automation solutions using Python, Shell, Ansible, or similar technologies.
Help eliminate manual operational activities through engineering and automation.
Improve monitoring, alerting, reporting, and operational workflows.
Drive incremental improvements that reduce risk, improve reliability, and increase operational efficiency.
Performance & Incident Management
·Analyze and resolve database performance issues.
Troubleshoot replication, backup/recovery, storage, network, and infrastructure-related incidents.
Participate in root cause analysis and drive permanent corrective actions.
Review operational metrics and trends to identify opportunities for improvement.
Operational Excellence
·Maintain accurate operational documentation, standards, and procedures.
Generate and present operational metrics, service health indicators, and reliability reporting.
Participate in incident response activities.
Demonstrate strong ownership from issue identification through resolution.
Required Qualifications- Strong experience administering enterprise database platforms, including:
- Sybase ASE
- Oracle RAC
- Additional database technologies such as MongoDB, Cassandra, Redis, PostgreSQL, MySQL, or similar platforms are a plus.
- Experience performing:
- o Installation
- Configuration
- Upgrades
- Patching
- Performance tuning
- Backup and recovery
- High availability and disaster recovery
- Experience with database replication technologies including:
- SAP Replication Server
- Data Guard
- HVR (preferred)
- Strong Linux administration skills.
- Experience with automation and scripting:
- Python
- Ansible
- Shell scripting
- Understanding of storage, networking, operating systems, and infrastructure services.
- Experience with Veritas Cluster Server, ASM, LVM, and SAN technologies.
- Familiarity with enterprise operational tooling such as Jira, Service Now and Confluence.
- Strong analytical, troubleshooting, and problem-solving skills.