F

Site Reliability Engineer

Fireblocks · 合规/托管/其他 · 上架 2026-05-28
面议
运维/SRE/基础设施中级📍 United States美国

职位要求 / 描述

<div class="content-intro"><p>The world of digital assets is accelerating in speed, magnitude, and complexity, opening the door to new ways for leveraging the blockchain. Fireblocks’ platform and network provide the simplest and most secure way for companies to work with digital assets and it trusted by some of the largest financial institutions, banks, globally-recognized brands, and Web3 companies in the world, including BNY Mellon, BNP Paribas, ANZ Bank, Revolut, and thousands more. </p></div><p><strong>About the team</strong></p> <p>Join a newly established, mission-critical SRE team at the forefront of financial infrastructure reliability. As part of Fireblocks Trust’s commitment to operational excellence, our Site Reliability Engineering team serves as the backbone of production systems, ensuring world-class uptime and performance for our digital asset custody and settlement platform.</p> <p><strong>What You'll Do</strong></p> <p>As part of your role, you would improve and establish new monitoring, alerting and observability of services using a wide range of tools. Additionally you would handle critical alerts and incidents and work directly with R&amp;D to improve and optimize availability.</p> <ul> <li>Research Fireblocks blockchain workflows, identify optimization opportunities, issues and improve monitoring.</li> <li>Help Identify root causes for incidents and prevent them from happening again. Solve and orchestrate outages by working with multiple teams.</li> <li>Improve and establish alerting for our infrastructure, services and business logic</li> <li>Work closely with the R&amp;D and Support: offering education and guidance on integration, support, and monitoring across the toolset</li> <li>Communicate and escalate issues to senior management in R&amp;D and support, write RCA’s, define next steps.</li> <li>Document actions in runbooks and then into automation using Python, Lamda, shell scripts, ArgoCD, Ansible.</li> <li>Focus on the system's observability, availability, reliability, performance/latency, monitoring</li> <li>Conduct periodic on-call duties and emergency response</li> </ul> <p><strong>What You’ll Bring:</strong></p> <p>To thrive in this role, you should bring the following qualifications and experience:</p> <ul> <li>At least 3+ years of experience as SRE, Infra Backend in a SaaS environment.</li> <li>You are curious, self-motivated, easy to work with, responsible and production aware. Fast learner and able to take a project from POC to production, while handling decision making and communication.</li> <li>Experience with Coding languages - Python/JavaScript/Bash (Must)</li> <li>At least 3+ years of experience with Alerting &amp; Monitoring systems such as <strong>DataDog</strong> Coralogix / Splunk / New Relic / Prometheus</li> <li>Experience working with Linux systems from kernel to shell and beyond</li> <li>Cloud systems such as <strong>AWS</strong> / Google cloud / Azure</li> <li>Configuration management such as <strong>Ansible</strong>/Chef/Puppet/ArgoCD</li> <li>Experience with Docker, Kubernetes and Helm</li> <li>SCM - Git/bitbucket/<strong>gitlab</strong>/Phabricator/gerrit</li> <li>High Analytical &amp; Troubleshooting skills - ability to solve complex problems</li> <li>Strong verbal and written communication skills and a collaborative mindset</li> </ul> <p><strong>Want to stand out of the crowd?</strong></p> <ul> <li>Previous experience in cryptocurrencies \ blockchains - big advantage</li> <li>In Depth knowledge in: Linux optimization, nginx, ArgoCD, DataDog, MySql</li> <li>Participated in Kubernetes migration projects</li> <li>Previous experience as C++ or Node developer</li> <li>BSC in Computer Science or related technical certifications</li> </ul> <div> <p>For employees hired to work remotely from New York, or from our NYC HQ, Fireblocks is required by law to include a reasonable estimate of the compensation range for this role. This range is specific to New York City and takes into con

技能关键字

#Python#JavaScript#C++#Validator/节点#Kubernetes#Docker#AWS#GCP

职责方向

稳定性保障故障/值班监控告警性能/容量自动化节点运维

数据来自公开渠道整理,薪资为公开 JD 或聚合估算,仅供参考,以面试谈薪为准。 ← 返回链聘 ChainHire 职位看板

Site Reliability Engineer · Fireblocks
面议
立即投递 →