Senior Application Support Engineer (SRE)
Oraclecloud · Tampa, FL, United States
About The Role
Are you ready to make an impact at DTCC?
Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.
Pay and Benefits
- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Pension / Retirement benefits
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
The impact you will have in this role
As a Senior Application Support Engineer (SRE), you will play a critical role in ensuring the stability, reliability, and performance of mission-critical applications at DTCC.
This role goes beyond traditional support—focusing on Site Reliability Engineering principles, proactive system improvement, and operational excellence. You will partner closely with development, infrastructure, and global operations teams to enhance system resilience, reduce operational toil, and drive continuous improvement across the platform.
Your Primary Responsibilities
- Act as a Lead Application Support Engineer with SRE responsibilities, partnering with engineering and infrastructure teams to improve system reliability, resilience, and observability
- Lead the resolution of critical production incidents, providing clear impact analysis, root cause identification, and preventive actions
- Own and drive incident, problem, and major incident management, including post-incident reviews and continuous improvement
- Proactively identify reliability risks and implement solutions to prevent recurrence and reduce operational toil
- Develop, maintain, and enhance runbooks, knowledge articles, and operational documentation
- Execute and support release, change, and deployment activities, including production releases and vendor upgrades
- Support and participate in Disaster Recovery (DR) testing, execution, and audit readiness
- Drive automation and alert optimization initiatives to improve efficiency and reduce noise
- Embed risk, control, and reliability best practices into day-to-day operations
- Collaborate with global teams to ensure high availability and operational excellence across systems
**NOTE: The Primary Responsibilities of this role are not limited to the details above. **
Qualifications
- 6+ years of experience in application support, SRE, or production engineering
- Bachelor's degree preferred or equivalent experience
Required Skills
- Strong hands-on experience in Application Production Support with SRE mindset focused on reliability, observability, production stability, resiliency, and incident prevention.
- Extensive experience supporting Linux and Windows environments, encompassing process inspection, log analysis, advanced troubleshooting, performance diagnostics, and system optimization.
- Strong scripting and automation expertise across Bash, Shell Scripting, Python, Ruby, Perl, and JavaScript.
- Hands-on experience with enterprise monitoring, observability, and analytics platforms such as Dynatrace, Splunk, Grafana, and Selenium.
- Deep technical expertise in SQL databases such as Oracle, Snowflake, PostgreSQL, or similar technologies, with proven ability to perform data analysis, troubleshoot production issues, conduct root cause investigations, and optimize query performance.
- Strong experience with ITSM and operational management platforms such as ServiceNow and Jira, supporting Incident, Problem, Change, and Major Incident Management processes.
- Hands-on experience supporting enterprise Messaging and Queueing Technologies such as IBM MQ, Oracle AQ, ActiveMQ, RabbitMQ, and Kafka.
- Strong experience working with enterprise job scheduling platforms, particularly Autosys.
- Working knowledge of Cloud and Container Technologies such as OpenShift, AWS (EC2, S3, Lambda, SQS, IAM Roles), Amazon RDS Aurora, PostgreSQL, and cloud-native operational support practices.
- Familiarity with Mainframe technologies and troubleshooting concepts across COBOL, JCL, DB2, DB2 Stored Procedures, CICS, SPUFI, and File-AID.
- Strong understanding of security, risk, and operational controls, including certificate management, password management, access controls, and security best practices.
- Exposure to Capital Markets and Financial Services environments with experience supporting mission-critical, high-availability, and highly regulated systems.
- Knowledge of Artificial Intelligence concepts and their practical application within Production Support, Operational Engineering, and Service Reliability disciplines.
- Demonstrated ability to communicate effectively across all organizational levels with exceptional verbal, written, and stakeholder management skills.
- Proven leadership qualities with a strong sense of ownership, accountability, urgency, and operational excellence in fast-paced production environments.
- Ability to collaborate effectively with global, geographically distributed teams, driving outcomes across technology, infrastructure, operations, and business functions.
- Proactive and continuous improvement mindset with a passion for automation, innovation, operational efficiency, reliability engineering, and reducing operational toil through engineering-led solutions.
- Strong analytical, problem-solving, and critical-thinking capabilities with the ability to navigate complex technical challenges, drive root-cause resolution, and deliver sustainable long-term solutions.
- Demonstrated ability to thrive in high-pressure environments while maintaining focus, composure, and an unwavering commitment to service availability, customer experience, and operational resilience.
The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Similar roles you might like
See all →This is an external listing. JobSpring does not represent or verify the employer. Report this listing
