| Job Position and Company | Posted | Location | Salary | Tags |
|---|---|---|---|---|
| $110k - $112k | ||||
Bitso📍 Latin America | $98k - $109k | |||
| $88k - $150k | ||||
| $80k - $101k | ||||
ISO 9001 Certified | 400+ students | Learn more | by Metana | ||
Okx📍 Remote | $140k - $144k | |||
| $96k - $192k | ||||
Alpaca📍 Remote | $119k - $135k | |||
| $140k - $200k | ||||
| $90k - $165k | ||||
| $119k - $170k | ||||
| $119k - $131k | ||||
| $105k - $120k | ||||
| $112k - $130k | ||||
| $103k - $120k | ||||
| $115k - $117k |
About Chainlink
Chainlink is the industry-standard oracle platform bringing the capital markets onchain and powering the majority of decentralized finance (DeFi). The Chainlink stack provides the essential data, interoperability, compliance, and privacy standards needed to power advanced blockchain use cases for institutional tokenized assets, lending, payments, stablecoins, and more. Since inventing decentralized oracle networks, Chainlink has enabled tens of trillions in transaction value and now secures the vast majority of DeFi.
Many of the world’s largest financial services institutions have also adopted Chainlink’s standards and infrastructure, including Swift, Euroclear, Mastercard, Fidelity International, UBS, S&P Dow Jones Indices, FTSE Russell, WisdomTree, ANZ, and top protocols such as Aave, Lido, GMX and many others. Chainlink leverages a novel fee model where offchain and onchain revenue from enterprise adoption is converted to LINK tokens and stored in a strategic Chainlink Reserve. Learn more at chain.link.
About the Role
As a Senior Site Reliability Engineer on the CCIP Platform team, you will ensure the reliability, scalability, and operational excellence of the systems powering Chainlink's Cross-Chain Interoperability Protocol (CCIP). This role exists to strengthen production resilience, reduce operational toil, and enable engineering teams to ship safely while maintaining high service availability. You will influence reliability practices across the platform and help establish operational standards that scale with the business.
Your Impact
Improve deployment safety and increase delivery velocity by advancing production engineering practices.
Establish distributed tracing across the platform to improve observability and accelerate incident investigation.
Eliminate operational toil through automation that increases engineering efficiency and platform reliability.
Drive adoption of meaningful SLOs, SLIs, and error budgets that guide engineering decisions and improve service health.
Increase platform scalability and operational readiness as CCIP continues to grow.
Strengthen Chainlink's reputation through highly available production systems while reducing operational overhead.
Requirements
Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems.
Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations.
Built and operated production Kubernetes environments supporting critical services.
Applied OpenTelemetry to improve observability across distributed systems.
Experience improving the reliability, scalability, and operability of production infrastructure.
Preferred Requirements
Demonstrated technical leadership influencing reliability practices across engineering teams.
Experience performing capacity planning and performance tuning for high-throughput distributed services.
Previous experience working on Web3 infrastructure or within a crypto-native engineering organization.
Applied chaos engineering or fault-injection techniques to improve production resilience.
Partnered with software engineering teams to conduct production-readiness reviews before service launches.
Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality.
All roles with Chainlink Labs are global and remote-based. Unless otherwise stated, we ask that you try to overlap some working hours with Eastern Standard Time (EST).
We carefully review all applications and aim to provide a response to every candidate within two weeks after the job posting closes. The closing date is listed on the job advert, so we encourage you to take the time to thoughtfully prepare your application. We want to fully consider your experience and skills, and you will hear from us regarding the status of your application shortly after the closing date.
Commitment to Equal Opportunity
Chainlink Labs is an equal opportunity employer. All qualified applicants will receive equal consideration for employment in compliance with applicable laws, regulations, or ordinances. If you need assistance or accommodation due to a disability or special need when applying for a role or in our recruitment process, please contact us via this form.
Global Data Privacy Notice for Job Candidates and Applicants
Information collected and processed as part of your Chainlink Labs Careers profile, and any job applications you choose to submit, is subject to our Recruiting Privacy Policy. By submitting your application, you are agreeing to our use and processing of your data as required.
What does Reliability Engineer do?
A Reliability Engineer is a professional who is responsible for ensuring the reliability and availability of systems and equipment in an organization
They use their knowledge of engineering principles, statistical analysis, and data science to identify and mitigate risks, prevent failures, and optimize system performance
Here are some of the typical tasks and responsibilities of a Reliability Engineer:
- Analyze data and perform statistical modeling: Reliability Engineers analyze data related to equipment performance, failure rates, and maintenance history to identify trends and patterns. They use statistical modeling to predict future failures and plan maintenance activities accordingly.
- Develop and implement reliability strategies: Reliability Engineers develop and implement strategies to improve the reliability and availability of equipment and systems. This may include performing root cause analysis, implementing preventive maintenance programs, and conducting failure mode and effects analysis (FMEA).
- Collaborate with other teams: Reliability Engineers collaborate with other teams such as operations, maintenance, and engineering to identify and address reliability issues. They may also work with suppliers to ensure the reliability of equipment and materials.
- Monitor and evaluate performance: Reliability Engineers monitor the performance of systems and equipment to identify areas for improvement. They use data to evaluate the effectiveness of reliability strategies and make adjustments as necessary.
- Provide technical support: Reliability Engineers provide technical support to other teams and stakeholders, answering questions and providing guidance on reliability-related issues.
- Continuously improve processes: Reliability Engineers are responsible for continuously improving reliability processes and methodologies. They stay up-to-date with the latest technologies and best practices in the field and identify opportunities for improvement.