Reliability Engineer

233 jobs found

Part of the Bondex Ecosystem

WxRK

Stop applying — get discovered by hiring agents.

Build your profile
Job Position and Company Posted Location Salary Tags

$77k - $85k

$210k - $310k

$72k - $72k

$98k - $112k

$180k - $218k

$140k - $200k

$211k - $249k

$122k - $140k

$103k - $117k

$170k - $210k

$186k - $218k

$90k - $145k

$90k - $145k

$90k - $145k

Zamp
$77k - $85k estimated
Bangalore

Site Reliability Engineer

Bangalore
Engineering – DevOps /
Full time /
On-site

Apply for this job
About Zamp:

At Zamp, we’re building AI agents that empower people to move at the speed of thought. Our vision is a world where AI handles the routine, so humans can focus on strategy and innovation. We are building a platform where all operational work runs autonomously. We partner with Fortune 500s, leading global banks and companies to streamline complex Finance and Operations processes.

Founded in 2022 by Amit Jain—an IIT Delhi and Stanford graduate with over 20 years of industry leadership, including roles as Managing Director at Sequoia Capital and Head of Asia Pacific at Uber—Zamp is backed by a stellar $22M seed round. Our investors include Sequoia Capital, Dara Khosrowshahi (CEO, Uber), Tony Xu (CEO, DoorDash), and other global visionaries.


About the team:

At Zamp, our engineering team is the force behind our technological innovations. Transforming the most ambitious ideas into reality, we breathe life into dreams through code and hardware. With the right mix of expertise and creativity, this team is the unseen magicians who add a dash of tech wizardry to make Zamp’s products shine.
From coding late into the night to brainstorming over a cup of coffee, we are always on a mission to make our technology stand out. We are not just the engineering team but the tech superheroes that keep Zamp at the forefront of innovation.

You are likely to succeed in this role if you bring experiences in :

    • Self-Hosted Infrastructure Ownership: Design, deploy, and maintain self-hosted systems and services, ensuring reliability, scalability, and security.
    • Kubernetes & Container Orchestration: Architect, manage, and scale Kubernetes clusters in production and self-hosted environments.
    • Infrastructure as Code (IaC): Write and manage Terraform modules to provision and manage infrastructure across AWS/GCP and on-prem setups.
    • CI/CD Automation: Build and maintain reliable CI/CD pipelines using tools like Jenkins, GitLab CI, or ArgoCD to ensure fast and safe deployments.
    • Monitoring & Observability: Set up and fine-tune observability tools like Grafana, Prometheus, and Graylog to monitor infrastructure, detect anomalies, and ensure uptime SLAs.
    • Scripting & Engineering: Write clean, modular automation scripts in Python, Bash, or Go to support operational needs and improve team productivity.
    • System Reliability & Incident Response: Own on-call responsibilities, drive root cause analysis, and continuously improve incident handling and system resilience.
    • Security & Compliance: Implement security and access controls within infrastructure, focusing on hardened self-hosted environments.

What we are actively looking for :

    • 4-6 years of experience
    • Proven experience managing self-hosted systems and internal tooling at scale
    • Deep hands-on knowledge of Kubernetes, including Helm, Ingress, scaling, and custom operators
    • Solid experience with Terraform for IaC; optionally Ansible for configuration
    • Expertise in monitoring/logging stacks: Grafana, Prometheus, Graylog, ELK
    • Hands-on with AWS, GCP, or Azure; strong understanding of cloud-native + on-prem hybrid setups
    • Strong scripting experience in Python, Bash, or Go
    • Proficiency in version control systems like Git for managing code repositories and facilitating collaboration among development teams
Our Culture and Benefits:

At Zamp, we promote a culture of open communication, collaboration, and empowerment. We value transparency, meritocracy, and a strong work ethic. Join our early team and help us build something exceptional.

Perks
- Competitive salaries and stock options with substantial potential upside.
- Collaborate with top talent.
- Diverse and inclusive workspace.
- Comprehensive medical insurance for employees, spouses, and children.
- A culture celebrating every victory.
- Continuous learning and skill development opportunities.
- Enjoy good food, games, and a comfortable office environment.

Apply for this job

What does Reliability Engineer do?

A Reliability Engineer is a professional who is responsible for ensuring the reliability and availability of systems and equipment in an organization

They use their knowledge of engineering principles, statistical analysis, and data science to identify and mitigate risks, prevent failures, and optimize system performance

Here are some of the typical tasks and responsibilities of a Reliability Engineer:

  1. Analyze data and perform statistical modeling: Reliability Engineers analyze data related to equipment performance, failure rates, and maintenance history to identify trends and patterns. They use statistical modeling to predict future failures and plan maintenance activities accordingly.
  2. Develop and implement reliability strategies: Reliability Engineers develop and implement strategies to improve the reliability and availability of equipment and systems. This may include performing root cause analysis, implementing preventive maintenance programs, and conducting failure mode and effects analysis (FMEA).
  3. Collaborate with other teams: Reliability Engineers collaborate with other teams such as operations, maintenance, and engineering to identify and address reliability issues. They may also work with suppliers to ensure the reliability of equipment and materials.
  4. Monitor and evaluate performance: Reliability Engineers monitor the performance of systems and equipment to identify areas for improvement. They use data to evaluate the effectiveness of reliability strategies and make adjustments as necessary.
  5. Provide technical support: Reliability Engineers provide technical support to other teams and stakeholders, answering questions and providing guidance on reliability-related issues.
  6. Continuously improve processes: Reliability Engineers are responsible for continuously improving reliability processes and methodologies. They stay up-to-date with the latest technologies and best practices in the field and identify opportunities for improvement.