RiseFlake
Acquirex

Senior Site Reliability Engineer (SRE) / DevOps Engineer

Acquirex

Pune, Maharashtra, India 10 yrs Posted 2 months ago 1 openings
Not disclosedIn OfficeCloud ArchitectDevOps EngineerSite Reliability Engineer

About this job

Are you an experienced and passionate Senior Site Reliability Engineer (SRE) or DevOps Engineer ready to tackle complex production challenges and drive operational excellence? Acquirex is actively seeking a highly skilled individual with a strong background in cloud infrastructure, automation, and observability to join our dynamic team in Pune. This pivotal role involves taking ownership of our production reliability, ensuring the stability and performance of our systems, and contributing significantly to our platform engineering initiatives.

As a Senior Site Reliability Engineer, you will be instrumental in building resilient, scalable, and secure cloud-native environments, applying your deep expertise to safeguard our critical services and empower our development teams. If you thrive in a challenging, fast-paced environment and possess a decade or more of hands-on experience in SRE or DevOps practices, we invite you to explore this opportunity to make a substantial impact.

Key Responsibilities

  • Lead and manage production incidents, providing 24/7 on-call support and driving efficient resolution.
  • Conduct thorough Root Cause Analysis (RCA) for all incidents and implement robust post-incident improvements to prevent recurrence.
  • Spearhead reliability engineering initiatives, focusing on enhancing system availability, reducing operational toil, and improving overall system resilience.
  • Administer and optimize our Kubernetes platform, overseeing containerized workloads and ensuring seamless deployment and operation.
  • Develop and maintain our infrastructure using Infrastructure as Code principles, primarily leveraging Terraform for cloud resource provisioning.
  • Implement and manage comprehensive monitoring, observability, and alerting solutions to ensure peak performance and proactive issue identification.
  • Design and execute disaster recovery plans, failover strategies, and system hardening procedures to ensure business continuity.
  • Uphold and enforce security best practices, ensuring compliance and proactively addressing vulnerability remediation across our infrastructure.
  • Collaborate closely with development teams to embed reliability from the ground up and foster a culture of shared ownership.

Requirements

  • 10+ years of demonstrable hands-on experience in Site Reliability Engineering (SRE) or DevOps Engineer roles.
  • Exceptional proficiency with Microsoft Azure, including deep knowledge of Azure Virtual Machines, Networking, Storage, Monitor, and Identity & Access Management.
  • Extensive experience with Kubernetes, Helm, and other container orchestration technologies.
  • Strong expertise in Infrastructure as Code (IaC) tools, particularly Terraform.
  • Solid background in Linux administration, networking fundamentals, and system troubleshooting.
  • Proficient in scripting languages such as Python and Bash for automation and operational tasks.
  • Hands-on experience with monitoring and observability platforms like Prometheus, Grafana, Datadog, Azure Monitor, or OpenTelemetry.
  • In-depth understanding of SLIs, SLOs, SLAs, Incident Management frameworks, and core Reliability Engineering principles.
  • Proven ability to troubleshoot complex distributed systems and make calm, effective decisions under pressure during critical incidents.
  • Experience with Git-based workflows using platforms like GitHub, GitLab, or Azure Repos.

What We Offer

  • A challenging and rewarding role at the forefront of cloud-native and reliability engineering.
  • Opportunity to work with cutting-edge technologies and drive significant impact on production systems.
  • A collaborative and supportive work environment that values innovation and continuous improvement.
  • Direct ownership over critical infrastructure and the chance to mentor and grow within a skilled team.
  • Competitive compensation and the opportunity to be an immediate joiner, making an impact from day one.

Eligibility

Professionals • 10-10 years of experience

Skills

AzureBashC#DatadogELK StackGitGrafanaHelmKubernetesLinuxOpenTelemetryPrometheusPythonTerraform

Perks & facilities

Certificate / Experience LetterLaptop / Equipment ProvidedStipend / SalaryTraining and Mentorship
Apply Now