Senior Software Engineer – Open Source, SWE-Bench Evaluation

🔥 18 hours ago

🇦🇷 Argentina – Remote

💵 $65 / hour

⏳ Contract/Temporary

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Anyone AI

Anyone AI

11 - 50 employees

Founded 2022

📚 Education

🤖 Artificial Intelligence

🎯 Recruiter

💰 $1.5M Pre Seed Round - Anyone AI on 2022-06

Education • Artificial Intelligence • Recruitment

Anyone AI is a Latin America–focused career accelerator and training academy that develops software engineers into AI and machine learning professionals through intensive, remote, mentor-led programs. It combines technical coursework (e. g. , Machine Learning Developer track), professional preparation (resume/LinkedIn optimization, interview coaching, English practice), and a placement-oriented model (income-share payments and access to employer partnerships) to help graduates secure higher-paying AI roles globally. Anyone AI also operates a talent marketplace connecting its community to companies seeking AI talent.

📋 Description

• Review ML challenges involving experiment design, model selection, datasets, metrics, preprocessing, distribution shift, contamination, label noise, feature leakage, hyperparameter tuning, train/validation/test methodology, reproducibility, and statistical significance • Determine whether challenges are technically sound, reproducible, appropriately difficult, and require strong machine learning reasoning • Evaluate whether datasets contain meaningful and learnable signals • Identify unintended shortcuts or artifacts in synthetic datasets • Determine whether tasks require genuine diagnosis of underlying ML problems rather than brute-force model selection or large hyperparameter searches • Review evaluation metrics and improvement thresholds • Detect metric gaming, data leakage, and evaluation flaws • Verify reproducibility across the complete data-to-model-to-evaluation pipeline • Assess whether challenge difficulty is appropriately calibrated • Provide recommendations for improving, recalibrating, or excluding problematic tasks • Analyze ML experiments, datasets, metrics, and pipelines for applied machine learning model-training and evaluation challenges

🎯 Requirements

• 3+ years of hands-on applied machine learning experience • Strong experience with ML experiment design, model selection, hyperparameter tuning, model evaluation, data preprocessing, and validation • Strong understanding of train, validation, and test splits • Ability to identify data leakage, label noise, distribution shift, spurious correlations, feature leakage, and data contamination • Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations • Strong understanding of ML evaluation metrics and when different metrics are appropriate • Experience debugging ML workloads across CPU and GPU environments • Ability to analyze technical problems and provide clear written feedback • Experience creating or participating in Kaggle, DrivenData, or similar ML competitions is nice to have • Experience designing benchmark datasets or ML challenges is nice to have • Background in data-centric AI or dataset quality is nice to have • Experience with synthetic data generation and validation is nice to have • Familiarity with statistical testing, confidence intervals, and effect sizes is nice to have • Experience with ML evaluation pipelines, RLHF, or AI model evaluation is nice to have • Experience developing ML curricula or technical assessments is nice to have • Understanding of shortcut learning, spurious correlations, Goodhart’s Law, Simpson’s paradox, and metric gaming is nice to have • Applicants must select a programming language or library for the interview and provide a location

🏖️ Benefits

• Remote work • Part-time, project-based consulting engagement

Apply Now

Similar Jobs

🕒 August 13

Coderio

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior Java Engineer designing and operating backend microservices for Coderio, a provider of scalable digital solutions. Applying AWS, Kafka, NoSQL, resilience patterns, and AI in complex fintech systems.

🗣️🇪🇸 Spanish Required

AWS

Java

Kafka

NoSQL

Go

🕒 July 30

AstroPay

201 - 500

💳 Fintech

👥 B2C

🤝 B2B

Product Lead / Senior Software Engineer at AstroPay, building the next generation of global financial infrastructure. Designing backend systems and scalable architectures within payments and fintech domains.

AWS

Cloud

Distributed Systems

Java

🕒 July 29

Ready

11 - 50

🎯 Recruiter

👥 HR Tech

🏪 Marketplace

Líder Técnico WSO2 para implementar plataformas de integración y middleware en proyectos tecnológicos de LATAM. Diseñando arquitecturas SOA, APIs y microservicios, y coordinando equipos técnicos e infraestructura.

🗣️🇪🇸 Spanish Required

Docker

Kubernetes

SOAP

🕒 July 23

Ryz Labs

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

Technical Lead guiding engineering for eDiscovery tools in LATAM, focusing on scalable product development. Leading a team to establish technical foundations and innovate in legal tech.

Cloud

🕒 June 29

Coderio

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

JavaScript

Kubernetes

Node.js

Python

React

SQL