Fusion HCR Is Hiring!
Role: Site Reliability Engineer
Location: Birmingham, Alabama
Work Arrangement: Remote
Employment Type: Full-time
Our client is an industrial marketplace serving customers across North America through an extensive network of locations and a broad portfolio of maintenance, repair, and operations products and services. The organization supports enterprise-scale operations across industrial automation, mechanical power transmission, electrical systems, hydraulics, industrial supplies, and material handling.
This opportunity is for a Site Reliability Engineer III who will help strengthen the reliability, scalability, and resilience of critical cloud platforms. The role partners closely with software engineering and infrastructure teams to improve automation, observability, performance, incident response, and cloud modernization. The ideal candidate brings strong experience supporting distributed production environments and is motivated by reducing operational effort while enabling rapid, reliable delivery.
Position Overview
Fusion HCR is seeking a Site Reliability Engineer III to bridge software engineering and systems administration across critical cloud platforms. This position will drive automation, improve platform stability, eliminate recurring manual work, and optimize performance across large-scale distributed environments.
The engineer will collaborate with development and infrastructure teams to balance feature velocity with dependable system operations. Responsibilities include managing cloud and containerized environments, strengthening observability, supporting incident response, troubleshooting complex infrastructure and application issues, and contributing to microservices and cloud modernization initiatives.
The successful candidate will use service-level objectives, monitoring data, logs, and performance metrics to identify risks before they become significant production issues. This role is well suited to an experienced SRE or DevOps professional who enjoys solving reliability problems through engineering, automation, and continuous improvement.
Key Responsibilities
-
Design, build, and maintain automated solutions that improve platform stability, capacity management, and deployment velocity.
-
Monitor production environments and use observability platforms to identify risks, troubleshoot performance bottlenecks, and support service-level objectives.
-
Respond to complex infrastructure, application, network, and platform incidents.
-
Participate in root-cause analysis and implement preventative improvements to minimize downtime and recurring failures.
-
Investigate malicious or anomalous traffic patterns and support mitigation efforts for enterprise applications.
-
Support Kubernetes platforms, containerized workloads, and elastic scaling across cloud environments.
-
Collaborate with software engineering teams on microservices architecture, platform administration, and cloud modernization.
-
Improve infrastructure-as-code practices and CI/CD processes.
-
Automate recurring operational tasks and strengthen alerting, monitoring, and deployment safeguards.
-
Partner with cross-functional teams to improve reliability, scalability, and production performance.
Required Qualifications
-
At least five years of experience in Site Reliability Engineering, DevOps, or Systems Engineering supporting enterprise-scale production environments.
-
Hands-on experience with Google Cloud Platform and Kubernetes, including containerization, cluster management, and elastic scaling.
-
Proven experience with Terraform and Azure DevOps or comparable continuous integration and continuous delivery tools.
-
Experience with observability and diagnostics platforms such as Dynatrace, Prometheus, and Grafana.
-
Strong knowledge of Linux and Windows administration.
-
Understanding of networking protocols, HTTP, proxies, and production network troubleshooting.
-
Experience supporting Java-based applications and microservices architectures.
-
Demonstrated experience with automation, incident response, root-cause analysis, and production reliability.
-
Ability to collaborate effectively with software development and infrastructure teams.
-
Authorized to work in the United States without current or future visa sponsorship.
-
Education in Computer Science, Information Technology, or a related field preferred.
Why Join Us
-
Opportunity to improve reliability across enterprise-scale cloud platforms.
-
Work on Kubernetes, automation, observability, infrastructure as code, and cloud modernization initiatives.
-
Collaborate with software engineering, infrastructure, and operations teams.
-
Help reduce manual operational work and prevent recurring production incidents.
-
Gain exposure to distributed systems, microservices, performance optimization, and traffic management.
-
Contribute to strategic technology initiatives within an established industrial organization.
-
Remote work arrangement with the opportunity to support critical platforms serving a broad customer base.
-
[N/A: Compensation, benefits, and formal career-development details were not provided.]