
Expert Site Reliability Engineer
Birlasoft Ltd
Hyderabad, Telangana, India 0-3 yrs Posted 2 months ago 1 openings
Not disclosedIn OfficeDevOps EngineerSite Reliability EngineerSoftware Development
About this job
Birlasoft is seeking a highly skilled and experienced Reliability Engineer (New Relic SME) to join our dynamic team across multiple locations in India. As a key member of our Site Reliability Engineering (SRE) function, you will be instrumental in ensuring the stability, performance, and scalability of our software systems. This pivotal Reliability Engineer role demands a proactive mindset, deep technical expertise in monitoring and automation, with a specific focus on Application Performance Monitoring (APM) using New Relic and other leading tools. If you are passionate about building robust, observable, and highly available systems, we encourage you to apply.
Key Responsibilities
- Proactively design, implement, and maintain effective monitoring systems that focus on symptom-based alerting to prevent incidents before they impact users.
- Utilize and optimize Application Performance Monitoring (APM) tools such as New Relic and Dynatrace to pinpoint bottlenecks, analyze application performance, and enhance resource utilization.
- Conduct in-depth log analysis using Splunk to identify anomalies, troubleshoot complex issues, and drive continuous improvements in system reliability.
- Develop and maintain informative dashboards and critical alerts to visualize system health, performance metrics, and ensure prompt incident response.
- Establish, track, and report on key reliability metrics, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Implement and champion observability practices, including distributed tracing, comprehensive logging, and robust metrics collection across our infrastructure.
- Collaborate closely with development, support, and operations teams to improve service quality through rigorous testing, streamlined release procedures, and post-incident reviews.
- Contribute to system design discussions and conduct capacity planning to ensure our infrastructure can scale efficiently to meet future demands.
- Perform root cause analysis for incidents, swiftly debug complex issues, and execute effective rollback procedures for faulty software deployments.
- Automate routine operational tasks and infrastructure management workflows using scripting languages like Python and Bash to enhance efficiency and reduce manual effort.
- Explore and integrate emerging technologies like Generative AI and Agentic AI frameworks to advance AIOps implementation across monitoring, observability, ITSM, and incident management.
- Provide mentorship and guidance to L1/L2 support teams, fostering best practices in monitoring, observability, and incident resolution.
- Manage and operate infrastructure using modern tools such as Chef, Ansible, Terraform, GitLab CI/CD, and Kubernetes, promoting an Infrastructure as Code (IaC) approach.
- Create and maintain clear, concise documentation for processes, procedures, and system configurations to promote knowledge sharing and reduce operational redundancy.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related field; or equivalent practical experience.
- Proven experience as a Site Reliability Engineer (SRE), DevOps Engineer, or a similar role focused on system reliability and performance.
- Demonstrated expertise as a Subject Matter Expert (SME) in Application Performance Monitoring (APM) tools, particularly New Relic, and experience with Dynatrace is a strong plus.
- Hands-on experience with log aggregation and analysis platforms like Splunk.
- Proficiency in automation scripting using Python and Bash.
- Solid understanding and practical experience with infrastructure management and orchestration tools such as Chef, Ansible, Terraform, and Kubernetes.
- Experience with Continuous Integration/Continuous Deployment (CI/CD) pipelines, ideally with GitLab CI/CD.
- Strong grasp of observability principles, including distributed tracing, metrics, and logging.
- Familiarity with AIOps concepts and a keen interest in leveraging AI for operational efficiency.
- Exceptional problem-solving, debugging, and incident management skills.
- Excellent communication and collaboration skills, with the ability to work effectively across cross-functional teams.
What We Offer
- An opportunity to work with cutting-edge technologies in a challenging and rewarding environment.
- A collaborative, innovative, and inclusive work culture that values continuous learning and professional growth.
- Engagement in high-impact projects that directly contribute to the stability and performance of critical systems.
- Competitive compensation and benefits package (to be discussed during the interview process).
- Clear pathways for career development and skill enhancement.
Eligibility
Professionals • 0-3 years of experience
Skills
AnsibleBashChefGitLab CI/CDKubernetesNew RelicPythonSplunkTerraform
Perks & facilities
Certificate / Experience LetterLaptop / Equipment ProvidedStipend / SalaryTraining and Mentorship