Software Engineer, Reliability

🕒 July 7

🇺🇸 United States – Remote

💵 $120k - $190k / year

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Nametag

Nametag

11 - 50 employees

Founded 2020

🏥 Healthcare

🛡️ Insurance

📦 Logistics

💰 Series unknown on 2021-02

Healthcare • Insurance • Logistics

Nametag is a workforce-focused identity verification company providing a SaaS platform that prevents deepfake impersonation and secures account recovery, onboarding, and helpdesk processes. Its Deepfake Defense™ engine combines cryptography, biometrics, and AI-resistant techniques to verify humans across channels, integrate with enterprise IAM and ITSM systems, and enable self-service MFA and password resets to reduce breaches and IT support costs.

📋 Description

• Design, build, and maintain scalable, cost-effective cloud infrastructure across AWS and Fly.io. • Own deployment strategy and service design across our monolith and microservices, including schema migrations and rollout safety. • Manage infrastructure-as-code (CloudFormation, Terraform or equivalent) for reproducible, auditable infrastructure. • Identify reliability risks, performance bottlenecks, and security gaps, and address them proactively. • Evolve our observability stack: logging, distributed tracing, metrics, and uptime monitoring. • Evolve on-call practices and incident response processes that keep us ahead of customer-impacting issues. • Champion a culture of reliability: postmortems, runbooks, and continuous improvement after incidents. • Manage and improve our CI/CD pipelines (GitHub Actions) and establish deployment best practices. • Build internal platform tooling and abstractions that reduce toil and increase engineering velocity. • Partner with product engineers to make infrastructure easy to use correctly and hard to use incorrectly. • Design and operate data pipelines that support our ML-powered verification systems. • Evolve our MLOps infrastructure so models can be trained, evaluated, and deployed safely and repeatedly. • Work closely with engineering and product leadership on technical roadmap decisions. • Review code, mentor peers, and help raise the bar on security, reliability, and operational discipline. • Communicate infrastructure tradeoffs clearly across technical and non-technical stakeholders.

🎯 Requirements

• Applicants must be legally authorized to work in the United States for any employer without current or future need for visa sponsorship. • Hands-on experience managing and securing cloud infrastructure. AWS required; Fly.io or similar a plus. • Production experience with Terraform or equivalent tools. • Deep experience with PostgreSQL, including schema design, migrations, and query performance tuning. • Strong experience designing and managing CI/CD pipelines, GitHub Actions preferred. • Proficiency in modern, type-safe languages. Go strongly preferred. • Experience building and operating logging, tracing, and metrics systems in production. • Prior experience shipping ML infrastructure into production, not just experimentation. • You've worked at an early-stage company and know what it means to move fast without compromising the things that matter. • Security-minded by default. You design with the threat model in mind, not as an afterthought.

🏖️ Benefits

• Competitive salary • Meaningful equity ownership • Comprehensive health benefits (medical, dental, vision) • Flexible paid time off • Quarterly team off-sites and travel support • New computer hardware and equipment • An inclusive environment where your voice has impact and your work drives change

Apply Now

Similar Jobs

🕒 July 7

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Lead Site Reliability Engineer overseeing reliability initiatives and SRE best practices at Empower. Architecting AWS infrastructure and mentoring SRE teams in a collaborative environment.

🕒 July 6

Infarsight

51 - 200

🤖 Artificial Intelligence

✈️ Travel

📦 Logistics

Senior DevOps & Cloud Infrastructure Engineer optimizing AWS environments for automation and product innovation. Leading deployment strategies and resource management across AWS, Vercel, and RackSpace.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 6

Waymark

11 - 50

📣 Marketing

🤖 Artificial Intelligence

Senior DevOps Engineer maintaining and improving AWS infrastructure at Waymark for Medicaid healthcare technology. Collaborating on system reliability and mentoring junior engineers.

🇺🇸 United States – Remote

💵 $109k - $194k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 July 6

Vytalize Health

201 - 500

🏥 Healthcare

☁️ SaaS

⚕️ Healthcare Insurance

DevSecOps Engineer responsible for implementing secure development and cloud security practices. Leading vulnerability management and integrating security within engineering and IT workflows.

🇺🇸 United States – Remote

💰 $100M Series C - Vytalize Health on 2023-02

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info

🕒 July 5

NVIDIA

10,000+ employees

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Senior Software Engineer at NVIDIA focusing on NVLink systems stability and reliability. Collaborating with multiple teams on innovative AI infrastructure solutions.

Distributed Systems

Python

Shell Scripting

Switching

TCP/IP