Production Support Engineer

🔥 6 hours ago

🌐 Argentina, Brazil, +1 more countries – Remote

infoinfo

⏰ Full Time

🟢 Junior

🟡 Mid-level

🏭 Production Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Solvd, Inc.

Solvd, Inc.

501 - 1000 employees

Founded 2010

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Solvd, Inc. is a global software engineering and consulting company that delivers comprehensive end-to-end solutions. Founded in 2011, Solvd offers core engineering services, including software development for web and mobile platforms, digital experience and design, data and AI/ML solutions, and cloud platform modernization. The company is an AWS partner and is known for optimizing resources, solving complex problems, and enhancing user satisfaction. Solvd emphasizes innovation and quality assurance, employing over 800 international professionals across 8 offices worldwide. The company prides itself on transforming businesses by providing feature-rich digital products, manual testing, automation, and employing cutting-edge technology solutions. Their proprietary tools like Zebrunner and Carina aid in improving development and QA processes. Solvd operates with a focus on collaboration, top talent cultivation, and customized technical solutions tailored to individual client needs, serving clients across 15 countries.

📋 Description

• Monitor platform health and triage incoming incidents • Distinguish critical issues such as outages and service degradation from non-critical bugs and defects • Investigate incidents using logging tools, API calls, responses, and log data • Own the incident management response process from first alert through post-mortem and corrective action follow-up • Notify stakeholders proactively about critical issues, SLA risk, and status via email, phone, or ticket system • Cross-reference tickets across multiple systems and follow defects through closure • Communicate technical findings in plain language to partners and call center teams • Manage the incident queue in Jira and prioritize bugs within engineering sprint cycles • Participate in weekly cross-functional meetings with engineering and account/call center management • Suggest continual improvements to applications and processes • Join on-call rotations after ramp-up and respond to OpsGenie alerts within defined SLA windows • Support the health and reliability of a revenue-critical sales platform between end users and engineering

🎯 Requirements

• 2+ years of troubleshooting and resolving issues for applications, servers, or infrastructure environments • 2+ years of providing clear status updates on tasks, issues, and resolutions to stakeholders at multiple levels • Working knowledge of APIs; able to read and interpret API calls and responses • Experience with Postman or similar API testing tools • Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic • Experience with SQL queries for troubleshooting and ad hoc reporting • Basic comfort reading HTML and JSON and using browser developer tools • Ability to participate in technical bridge calls and follow incidents through to resolution • Exceptional communication skills across technical and non-technical audiences • Strong time management, prioritization, and organizational skills under pressure • Customer service mindset and genuine interest in supporting end users • Empathy, humility, and comfort with ambiguity • Available during U.S. Eastern business hours (9 AM–6 PM ET); Eastern timezone strongly preferred for onboarding and on-call coordination • Bachelor's degree in a related field or equivalent work experience • Preferred: AWS Cloud Practitioner level or above • Preferred: Familiarity with Git • Preferred: Terraform or similar infrastructure-as-code concepts • Preferred: OpsGenie or similar alerting platforms • Preferred: AI tooling for troubleshooting and investigation workflows • Preferred: Engineering deployment lifecycle and release processes • Preferred: On-call rotation structures and incident severity frameworks • Preferred: Partner or call center communication management during live incidents • Preferred: Travel, hospitality, or high-volume transactional platform experience

🏖️ Benefits

• Comprehensive onboarding documentation and a structured 6-month ramp to full self-sufficiency • On-call rotations begin only when you're ready, with manager backup during early rotations • Active alert window is 8 AM–1 AM Eastern; overnight suppression windows are built in • Global team collaboration across continents and cultures • Inclusive environment prioritizing continuous learning, innovation, and ethical AI standards

Apply Now