SRE Monitoring Platform Software Engineer – Early Career, Temporary

🔥 0 minutes ago

🏄 California, Texas – Remote

infoinfo

💵 $105k - $155k / year

⏰ Full Time

🟢 Junior

🟡 Mid-level

🧑‍💻 Full-stack Engineer

🚫👨‍🎓 No degree required

👻 Ghost score 5%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Bitdeer Group

Bitdeer Group

201 - 500 employees

💼 Consulting

📦 Logistics

🏗️ Construction

💰 Post-IPO Equity on 2023-05

Consulting • Logistics • Construction

Bitdeer Group is a leader in the blockchain and high-performance computing industry. It is one of the world’s largest holders and suppliers of hash rate, offering specialized mining infrastructure and high-quality hash rate sharing products. Founded by cryptocurrency pioneer Jihan Wu and led by CEO Matt Linghui Kong, the company is headquartered in Singapore with mining datacenters in the United States, Norway, and Bhutan. Bitdeer is committed to providing comprehensive computing solutions, including cloud services and AI capabilities, while emphasizing dedication, authenticity, and trustworthiness in its mission to be the most reliable provider in the industry.

📋 Description

• Contribute to the NeoCloud SRE platform, a multi-region system operating a GPU rental fleet across data centers • Build well-scoped components from design through production code under senior-engineer guidance • Ship code through GitOps and CI/CD release pipelines while following Plugin Framework conventions and declared SLOs • Build collection and storage components, including ingestion, query, storage-path, enrichment, and collection-monitor code • Contribute to alerting, correlation, and SLO frameworks; implement and tune default alert rules • Develop topology and cluster-health services and collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay • Help build remediation actuators, orchestration/workflow components, inspection probes, and job schedulers • Instrument services with metrics, logs, and traces using OpenTelemetry • Build dashboards and write actionable on-call runbooks • Write unit, integration, and contract tests for shipped work • Participate in chaos and soak tests led by senior engineers • Operate built systems under guidance and participate in on-call as a shadow before taking primary responsibility • Progress toward independently delivering components and owning a sub-context within 12 months

🎯 Requirements

• 0–2 years of software engineering experience; new graduates with strong projects or internships welcome • Solid fundamentals in one programming language: Go (preferred), Python, Java, or Rust • Ability to write clean, tested, readable code and explain design choices • Knowledge of data structures, algorithms, concurrency, TCP/HTTP networking, and operating-system concepts • Understanding of distributed-systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency • Hands-on exposure to monitoring/observability tools such as Prometheus, Grafana, or Loki; ability to write basic PromQL and instrument a service • Familiarity with Linux, shell, system logs, and standard debugging tools • Kubernetes fundamentals, including Pods, Services, and Deployments; experience running something on Kubernetes • Git and CI fundamentals, including branching, pull requests, and use of a CI pipeline • Habit of writing unit and integration tests • Clear written and verbal English • Curiosity and eagerness to learn GPU/AI infrastructure, AIOps, distributed systems, and observability • Nice-to-have: internship or project in monitoring/observability, telemetry pipelines, or platform/SRE tooling • Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray • Nice-to-have: exposure to AIOps/ML-adjacent tooling • Nice-to-have: contributions to open-source observability or cloud-native projects

🏖️ Benefits

• Mentorship from senior and principal engineers • Guided participation in on-call, initially as a shadow before taking primary responsibility • Exposure to production-scale observability, GPU/AI infrastructure, AIOps, distributed systems, and cloud-native technologies • Opportunity to contribute to greenfield systems with established Plugin Framework, GitOps pipeline, and SLO framework • Equal employment opportunities

Apply Now

Similar Jobs

🔥 1 hour ago

Walmart

10,000+ employees

🛒 Retail

🛍️ eCommerce

👥 B2C

Software Engineer III building scalable data and AI platform solutions at Walmart. Leading projects, production troubleshooting, automation, and engineering team mentorship.

🇺🇸 United States – Remote

💵 $108k - $216k / year

💰 $5G Post-IPO Debt - Walmart on 2023-04

⏰ Full Time

🟢 Junior

🟡 Mid-level

🧑‍💻 Full-stack Engineer

🔥 3 hours ago

Nous Research

11 - 50

🤖 Artificial Intelligence

🔬 Science

Mobile Software Engineer building Hermes Agent’s iOS and Android apps. Owning native, cross-platform, API, testing, and AI-agent mobile experiences.

🔥 3 hours ago

ActioNet, Inc.

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Software Engineer building Java, Python, SQL, and Cube.js solutions for ActioNet’s data-driven applications and analytics platforms. Developing scalable backend services, APIs, and semantic data models.

🔥 4 hours ago

ActioNet, Inc.

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Full Stack Developer modernizing data-driven applications and analytics platforms for ActioNet, an IT solutions provider. Building Java, Python, SQL, and Cube.js solutions.

🔥 5 hours ago

Netflix

10,000+ employees

📱 Media

👥 B2C

Group Tech Lead shaping Netflix’s application networking strategy and architecture. Securing and scaling the communication backbone powering Netflix streaming, Live, Ads, and Gaming.