Site Reliability Engineer – Low-Latency Trading Systems

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Ondo Finance

Ondo Finance

51 - 200 employees

₿ Crypto

💳 Fintech

💸 Finance

💰 Initial Coin Offering - Ondo Finance on 2024-01

Crypto • Fintech • Finance

Ondo is a company building institutional-grade onchain financial infrastructure and products that bridge traditional finance (TradFi) and decentralized finance (DeFi). Its offerings include tokenized public securities (Ondo Stocks) that make public securities freely transferable and usable in DeFi, USDY (a permissionless yield-bearing stablecoin), and OUSG (an institutional product offering exposure to short-term US Treasuries with 24/7 instant minting and redemptions). Ondo emphasizes compliance, institutional-grade security, third-party audits, and partnerships with asset managers and regulated service providers. The company’s Nexus technology enables instant minting and redemption for tokenized US Treasuries and stablecoins and supports omnichain issuance and distribution. Ondo works with partners like Broadridge, J. P. Morgan, Mastercard, Ripple, and asset managers to bring voting and cross-border redemption capabilities to tokenized assets.

📋 Description

• Own production reliability for real-time trading services, including trading engines, execution gateways, market data ingestion, and PnL/reconciliation pipelines • Operate and evolve multi-region Kubernetes clusters on AWS EKS using GitOps with Flux and SOPS-encrypted secrets • Build and refine Prometheus metrics and alerting, Datadog logs and dashboards, and SLOs • Improve deployment safety through progressive rollouts, configuration reload behavior, and safeguards for live trading • Debug production incidents involving stale market data, exchange rate limits, WebSocket disconnects, order-lifecycle desynchronization, and trading-path latency regressions • Harden market data ingestion from Databento and venue-native REST/WebSocket feeds through staleness detection, failover, and replay • Build reconciliation and data-integrity tooling across live gauges, Postgres, and S3 Parquet data lake • Participate in on-call rotation covering US equity market hours and 24/7 crypto venues

🎯 Requirements

• 5+ years in SRE, production engineering, or infrastructure roles, with meaningful time supporting real-time or latency-sensitive systems • Strong programming ability in Go or Rust, and willingness to work in both • Deep, hands-on Kubernetes and AWS experience running stateful, latency-sensitive workloads in production • Fluent PromQL, structured-log analysis, and experience designing high-signal, low-noise alerts • Solid Linux internals and networking fundamentals • Ability to chase p99 regressions through the kernel, NIC, or GC • Experience participating in incident response and communicating clearly during and after incidents • Must be based in the United States

🏖️ Benefits

• Competitive compensation including future token rights and/or equity according to preferences • Full medical, vision, and dental benefits • Flexible vacation policy (PTO) • Remote-first work arrangement • Opportunity to help shape the company’s vision, culture, and design practices • Collaboration with A+ colleagues and leading industry experts

Apply Now

Similar Jobs

🔥 13 minutes ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer operating Peraton’s AWS and OpenShift production infrastructure. Managing reliability, observability, incident response, releases, and infrastructure automation for national security missions.

🔥 13 minutes ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer operating Peraton’s AWS, GovCloud, and ROSA production infrastructure. Improving observability, incident response, resilience, and deployment automation for national security systems.

🔥 13 minutes ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Site Reliability Engineer operating AWS, Azure, and GCP production infrastructure for Peraton, a national security and enterprise IT provider. Improving reliability, observability, incident response, and deployment automation.

🔥 3 hours ago

Gifthealth

501 - 1000

🏥 Healthcare

📦 Logistics

💼 Consulting

DevSecOps Engineer embedding automated security across Gifthealth’s prescription healthcare platform. Building CI/CD, application, infrastructure, container, and Kubernetes security controls.

🇺🇸 United States – Remote

💵 $115k - $165k / year

💰 $40M Private Equity Round - GiftHealth on 2023-04

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 8 hours ago

Replit

51 - 200

🤖 Artificial Intelligence

🤝 B2B

Senior Site Reliability Engineer ensuring Replit’s reliable, scalable infrastructure serving millions of developers. Automating operations, observability, incident response, and performance optimization.