C

Senior Site Reliability Engineer (L3)

CoinGecko · 节点/RPC基础设施 · 上架 2026-06-01
面议
运维/SRE/基础设施资深/Staff📍 Malaysia远程

职位要求 / 描述

CoinGecko is a global leader in tracking cryptocurrency data. Operating since 2014, CoinGecko has built the world's largest cryptocurrency data platform, tracking over 10,000 tokens across more than 400 exchanges, serving over 300 million page views in more than 100 countries, and handling more than 45 billion API calls every month. We are proud to have played a major part in mainstream awareness, adoption, and education of cryptocurrency globally. We at CoinGecko believe that cryptocurrency and blockchain will define the future of finance, bringing greater financial and economic freedom around the world. In anticipation of that future, CoinGecko is building the foundation to scale cryptocurrency market data to serve billions. We practice transparent salaries and a level structure at CoinGecko: • The salary range for the L3 position is RM16,899 - RM18,589. • For more junior candidates, we may evaluate you as a L2, with a salary range of RM12,053 - RM13,258. • Learn more about our level structure at CoinGecko's Career Progression.    Note: This is a remote-first role based in Malaysia, and only candidates based in Malaysia will be considered. Job Responsibilities: • System Architecture: Review architecture and software components with software engineers. Ensure best practices are consistent across all teams. • Operational Excellence: Own and ensure SLOs and SLAs are met. Monitor operational metrics and lead improvement plans. Develop and maintain tools including infra-as-code resources to scale operations and allow other teams to be autonomous. • Security and Compliance: Manage and audit security controls to meet enterprise requirements. Implement and maintain best practices and compliance standards. Collaborate with legal and compliance to assess overall risk management. • Release Planning: Lead strategic release plans (e.g., canary or blue-green deployments) to reduce blast radius and allow for faster reversal during release failures. Work closely with developers for pre-release requirements including provisioning test environments. Conduct ad hoc performance tests based on requirements. • Incident Management: Lead incident response and post-mortems to resolve production issues, identify root-causes and prevent future occurrences. • Disaster Recovery: Develop and implement DR plans and procedures, including data recovery and fault injection simulations on production replica. • Daily Operations: Perform and improve day-to-day tasks including access onboarding-offboarding, config and patch management etc. Plan capacity to ensure our systems have sufficient capacity to handle peak demand while optimizing cost. • Documentation: Develop and extend runbooks, documentation and other technical assets. Support periodic technical audits as required. • Sharpen the Saw: Stay up-to-date with emerging trends and technologies in software development and contribute to knowledge sharing. Learn advanced architecture standards and new tools that improve the team’s code base and productivity. Demonstrate thorough understanding of a subject matter and how to apply it effectively. • Team player: Collaborating with cross-functional teams to ensure smooth deployment and operation of software releases. Answer technical questions from other teams or outside the organization. • Coaching: Provide feedback on the performance of junior staff and participate in people development initiatives. • Support any ad hoc tasks as required by the company. Job Requirements: • Proven track record: 3 to 5 years in managing software deployments and instrumentation in production environments with defined SLAs and SLOs. Strong knowledge of software delivery and devops principles. • Cloud Operations: Experience with cloud platforms (e.g., AWS, CloudFlare, GCP) and infrastructure-as-code tools (e.g., Terraform, CloudFormation). Strong programming and scripting skills, preferably in languages such as Python, Go, or Ruby. • Accreditation: Bachelor’s degree

技能关键字

#安全审计#合规

职责方向

架构设计稳定性保障故障/值班监控告警性能/容量部署发布

数据来自公开渠道整理,薪资为公开 JD 或聚合估算,仅供参考,以面试谈薪为准。 ← 返回链聘 ChainHire 职位看板

Senior Site Reliability Engineer (L3) · CoinGecko
面议
立即投递 →