Job Description: DevOps / Site Reliability Engineer (SRE)
Company Overview: NextAmp LLC is a digital modernization services company focused on transforming the insurance industry. We help insurance organizations redesign their workflows and build AI-powered, automated systems that improve efficiency and accuracy. Our solutions streamline key operational areas such as claims processing, underwriting, billing, and back-office operations.
Position: DevOps / Site Reliability Engineer (SRE)
Location: USA (Remote)
Experience level: 5+ years
About the Role
We are looking for an experienced DevOps / Site Reliability Engineer (SRE) with strong expertise in cloud infrastructure, automation, observability, and production operations. The ideal candidate will have hands-on experience with Python, Apache Airflow, Astronomer, GitHub Actions, Amazon EKS, and the Grafana observability stack. This role involves building scalable infrastructure, improving CI/CD pipelines, implementing modern monitoring solutions, and ensuring the reliability and performance of mission-critical applications.
Key Responsibilities
- Design, implement, and maintain scalable and secure cloud infrastructure on AWS.
- Build, automate, and optimize CI/CD pipelines using GitHub Actions and other modern DevOps tools.
- Lead migration of Apache Airflow v2 to Airflow v3 on the Astronomer platform.
- Design, develop, optimize, and troubleshoot Airflow DAGs for production data workflows.
- Deploy, manage, and maintain Kubernetes workloads on Amazon EKS.
- Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or similar tools.
- Implement and maintain observability solutions using the Grafana Stack, including Grafana Alloy and Grafana Beyla.
- Utilize eBPF technologies for deep infrastructure and application observability.
- Build and maintain automated testing infrastructure supporting CI/CD pipelines.
- Monitor application and infrastructure health, respond to production incidents, and perform Root Cause Analysis (RCA).
- Automate operational processes to improve deployment reliability and reduce manual effort.
- Collaborate closely with software engineering teams to improve deployment processes, platform stability, and system reliability.
- Implement logging, monitoring, alerting, and performance optimization best practices.
- Ensure high availability, scalability, security, and operational excellence across production environments.
Required Skills
- 5–8 years of experience in DevOps or Site Reliability Engineering.
- Strong programming experience in Python.
- Hands-on experience migrating Apache Airflow v2 to Airflow v3.
- Experience working with the Astronomer platform.
- Strong knowledge of Airflow DAG development, scheduling, monitoring, debugging, and optimization.
- Experience building and managing CI/CD pipelines using GitHub Actions.
- Experience designing and maintaining CI testing infrastructure.
- Strong experience with Docker and Kubernetes.
- Hands-on experience managing workloads on Amazon EKS (Elastic Kubernetes Service).
- Strong knowledge of Infrastructure as Code using Terraform, CloudFormation, or similar tools.
- Experience with AWS cloud services and cloud-native architectures.
- Hands-on experience with the Grafana Stack, including:
- Grafana
- Grafana Alloy
- Grafana Beyla
- Deep understanding of eBPF for Linux observability, networking, and performance monitoring.
- Experience implementing monitoring, logging, tracing, and alerting solutions.
- Strong Linux administration, networking, and security fundamentals.
- Strong troubleshooting skills with production incident management experience.
Preferred Skills
- Experience with Helm, ArgoCD, or FluxCD.
- Experience with OpenTelemetry and distributed tracing.
- Knowledge of Prometheus, Loki, Tempo, and modern observability platforms.
- Experience with service mesh technologies such as Istio or Linkerd.
- Familiarity with secrets management solutions such as HashiCorp Vault or AWS Secrets Manager.
- Knowledge of SRE principles including SLIs, SLOs, Error Budgets, and reliability engineering.
- Experience supporting highly available, distributed production systems.
- Experience implementing cloud governance and infrastructure cost optimization.
- Knowledge of disaster recovery, backup, and business continuity strategies.
Nice to Have
- Experience supporting 24×7 production environments.
- Experience working with large-scale distributed platforms and cloud-native applications.
- AWS Certified DevOps Engineer.
- AWS Solutions Architect.
- Certified Kubernetes Administrator (CKA).
- Grafana or Kubernetes-related certifications.
- Microsoft Azure DevOps Engineer Expert (if applicable).
Working Hours: 8am EST – 5pm EST
Pay: $35.00 - $49.00 per hour
Application Question(s):
- Work Authorization Status *
Please specify your current work authorization (e.g., U.S. Citizen, Green Card, H-1B, H-4 EAD, F-1 OPT/STEM OPT, TN Visa, etc.).
- Notice Period (in days)
- Salary Expectation
Work Location: Remote