Search Remote Jobs

Site Reliability Engineer

Job not on LinkedIn

🔥 0 minutes ago

🏄 California, New York, +1 more states – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

đźź  Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

đź‘» Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Databento

Databento

11 - 50 employees

Founded 2019

đź’Ľ Consulting

📣 Marketing

📦 Logistics

đź’° $24.3M Series A on 2021-11

Consulting • Marketing • Logistics

Databento is a company that provides comprehensive market data solutions through APIs for both real-time and historical data. Founded in 2018, it delivers normalized data from a wide array of asset classes including futures, options, and equities, sourced directly from colocation sites to ensure low latency and high reliability. Databento offers a simplified and efficient way to access market data, with customizable pricing models and extensive support for developers. Their services include live data streaming, historical data retrieval, and detailed insights into corporate actions. The platform caters to over 3,000 leading firms and startups, utilizing advanced technology to offer solutions such as full exchange order book replay and seamless integration with programming languages like Python and C++.

đź“‹ Description

• Own uptime, SLAs, and SLOs across API and platform services • Set reliability and operational best practices for developers • Build and maintain observability across logging, metrics, and tracing • Design and run high-availability deployment and containerization strategies • Profile and optimize Python applications for throughput, latency, and cost • Debug production issues down to the OS level using strace, perf, eBPF, ss, and gdb • Improve deployment and CI/CD workflows • Participate in the on-call rotation, lead incident response, and run post-incident reviews • Identify needed fixes and take projects from idea to completion • Work on petabyte-scale data processing, customer management and billing, and query systems powering APIs

🎯 Requirements

• Midlevel or senior individual contributor • Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading firm, tech company, or high-growth startup • Hands-on experience with observability tooling for logging, metrics, and tracing, such as Prometheus, OpenTelemetry, VictoriaMetrics, Jaeger, Logstash, Loki, or Vector • Experience with containerization and high-availability deployment, such as Docker, Podman, Docker Compose, Docker Swarm, Kubernetes, or k3s • Strong proficiency in Python, including application development and performance optimization • Comfortable with Linux debugging and profiling tools such as strace, perf, eBPF, ss, and gdb • Track record of measurable impact in a recent role • Experience with alerting and incident response best practices is a plus • Familiarity with configuration management or infrastructure-as-code tools such as Ansible or Terraform is helpful • HTTP benchmarking, load testing, and capacity planning experience is a bonus • Database schema design and query optimization skills are nice to have • Good communication skills and work ethic for a remote workplace • Interest in financial data or algorithmic trading

🏖️ Benefits

• Equal employment opportunities and nondiscrimination protections • Employment accommodation available upon request • AI-powered Talent Matching opt-out option

Apply Now

Similar Jobs

🔥 30 minutes ago

Peraton

10,000+ employees

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

Senior SRE building Python, AWS, and Terraform reliability solutions for Peraton’s national security missions. Improving cloud platform reliability through automation, observability, and incident management.

🔥 8 hours ago

The SSI Group

11 - 50

🎯 Recruiter

👥 HR Tech

🤝 B2B

DevOps Process Analyst optimizing development and operations workflows for SSI Group, a healthcare software company. Improving processes, reporting, risk management, and cross-functional delivery.

🔥 9 hours ago

CentralReach

201 - 500

🏥 Healthcare

📚 Education

Site Reliability Engineer operating AWS and Snowflake infrastructure for CentralReach’s autism and IDD care software. Improving reliability, connectivity, deployments, and incident response across data platforms.

🔥 9 hours ago

CentralReach

201 - 500

🏥 Healthcare

📚 Education

Senior SRE improving reliability, observability, and cloud infrastructure for CentralReach’s autism and IDD care software platforms. Driving SLO adoption, incident response, automation, and production performance.

🔥 10 hours ago

Katmai Government Services

1001 - 5000

🏛️ Government

đź“‹ Compliance

⚕️ Healthcare Insurance

DevOps Engineer II automating CI/CD, containers, and infrastructure for Navy readiness and training systems. Supporting DoD cybersecurity compliance across development, test, and production environments.