Overview
Salary: $48.48-53.87 Hourly
Join a leading innovator in the financial services industry, a company dedicated to empowering individuals and institutions to achieve their financial goals. This organization stands at the forefront of technological advancement, continuously evolving its platforms and services to deliver unparalleled value and security to millions of clients. Aquent is proud to partner with this client, offering you a unique opportunity to contribute to their mission-critical operations and drive significant impact. Are you a highly motivated and experienced reliability professional passionate about automation and operational excellence? This is your chance to join a pivotal team where you will be instrumental in shaping the future of our client's infrastructure. In this dynamic role, you will elevate the reliability and operational excellence of critical production systems, driving automation and ensuring seamless operations across both on-premises and cutting-edge cloud platforms. Your expertise will directly impact the stability, performance, and scalability of systems that serve millions, making a tangible difference in the daily lives of clients. If you thrive on solving complex challenges, possess an automation-first mindset, and are eager to contribute to a high-impact environment, we invite you to shine with us! What You'll Do
- Develop Python-based automation solutions to reduce manual operational effort and enhance efficiency across diverse environments.
- Automate infrastructure management across Linux, Windows, Kubernetes, cloud platforms, and cloud-native environments.
- Integrate various tools and platforms through APIs and client libraries to streamline workflows and improve system interoperability.
- Assist in implementing robust infrastructure automation using industry-standard technologies to build scalable and resilient systems.
- Support CI/CD automation and deployment reliability initiatives to ensure smooth, consistent, and secure software releases.
- Monitor and maintain production systems to consistently meet and exceed reliability and availability objectives, ensuring seamless service delivery.
- Actively participate in incident response, thorough troubleshooting, and root cause analysis activities to prevent recurrence and enhance system stability.
- Develop proactive automation and operational improvements to eliminate recurring issues and continuously improve system resilience.
- Support disaster recovery, failover testing, and other operational readiness activities to ensure business continuity and minimize downtime.
- Perform in-depth performance analysis and regular system health reviews to optimize system behavior and identify areas for improvement.
- Build and maintain comprehensive dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, or similar industry-leading tools.
- Improve visibility into application and infrastructure health through enhanced metrics, logs, and traces for proactive issue detection.
- Investigate alerts thoroughly and identify opportunities to reduce noise and improve detection accuracy, enhancing operational responsiveness.
- Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics to advance operational capabilities.
- Assist in developing automation solutions that leverage AI to significantly improve operational efficiency and predictive insights.
- Participate in evaluating emerging AIOps capabilities and cutting-edge observability technologies to keep our systems at the forefront of innovation.
What You'll Bring
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3 to 5 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, Systems Engineering, or Platform Engineering, with a strong focus on production operations.
- Must have supported production systems at scale, demonstrating a deep understanding of operational challenges in large-scale environments.
- Proven experience in operations ownership, incident response, and reliability engineering.
- Strong programming skills in Python for automation and tooling development; Python must be a primary skill, not just basic scripting.
- Demonstrated experience building automation tools, scripts, frameworks, or operational solutions, with the ability to provide examples of personally developed automation.
- Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
- Familiarity with infrastructure automation and configuration management tools.
- Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
- Solid understanding of Linux systems, networking, and distributed applications.
- Strong analytical, troubleshooting, and problem-solving skills, with experience in critical production incident management and Root Cause Analysis (RCA).
- Ability to work effectively in fast-paced, mission-critical environments, consistently reducing operational toil through automation.
What Will Set You Apart
- Experience with Terraform, Ansible, or other Infrastructure as Code solutions.
- Exposure to OpenTelemetry and modern observability practices.
- Experience with CI/CD pipelines and deployment automation.
- Knowledge of AI/ML, AIOps, or intelligent operational tooling.
- Experience supporting highly available production systems in regulated or enterprise environments.
About Aquent Talent Aquent Talent connects the best talent in marketing, creative, and design with the world's biggest brands. Our eligible talent get access to amazing benefits like subsidized health, vision, and dental plans, paid sick leave, and retirement plans with a match. We also offer free online training through Aquent Gymnasium. More information on our awesome benefits! Aquent is an equal-opportunity employer. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, and other legally protected characteristics. We're about creating an inclusive environment-one where different backgrounds, experiences, and perspectives are valued, and everyone can contribute, grow their careers, and thrive. #LI-LP1
|