Member of Engineering – Pre-training, Synthetic Data

🕒 January 29

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🖥 Software Engineer

👻 Ghost score 41%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of poolside

poolside

51 - 200 employees

Founded 2023

🤖 Artificial Intelligence

🏢 Enterprise

Artificial Intelligence • Enterprise

poolside is a frontier AI lab and enterprise platform that builds and deploys foundation models, multi-agent systems, and developer-facing tools focused on automating complex software work. The company specializes in on-prem and VPC deployments, security-first integrations, governance, and connectors to enterprise data sources so organizations can run agents and models inside their own boundaries. Poolside embeds research and engineering with customers to deliver outcome ownership, risk controls, and measurable business impact while advancing toward AGI by starting in high-consequence software environments.

📋 Description

• You’ll be working on our data team focused on the quality of the datasets being delivered for training our models. • This is a hands-on role where your #1 mission would be to improve the quality of the pretraining datasets by leveraging your previous experience, intuition and training experiments. • This role particularly focuses on generating synthetic data at scale and determining the best strategies to leverage such data into training large models. • You’ll closely collaborate with other teams like Pretraining, Postraining, Evals, and Product to define high-quality data needs that map to missing model capabilities and downstream use cases. • Staying in sync with the latest research in synthetic data generation and pretraining is key to success in this role. • You will constantly lead original research initiatives through short, time-bounded experiments while deploying highly technical engineering solutions into production. • With the volumes of data to process being massive, you'll have a performant distributed data pipeline together with a large GPU cluster at your disposal. • To deliver large, high-quality, and diverse synthetic datasets mixing natural language and code modalities to train best-in-class coding agents.

🎯 Requirements

• Strong machine learning and engineering background • Experience with Large Language Models (LLM) • Understanding of how LLMs learn • Data ablations and scaling laws • Post-training techniques • Training reasoning and agentic models • Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc. • Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc) • Experience in building trillion-scale pretraining datasets, and familiarity with concepts like data curation, deduplication, data mixing, tokenization, curriculum, impact of data repetition, etc. • Excellent programming skills in Python • Strong prompt engineering skills • Experience working with large-scale GPU clusters and distributed data pipelines • Strong obsession with data quality • Research experience: Author of scientific papers on any of the topics: applied deep learning, LLMs, source code generation, etc. - is a nice to have • Can freely discuss the latest papers and descend to fine details • Is reasonably opinionated

🏖️ Benefits

• Fully remote work & flexible hours • 37 days/year of vacation & holidays • Health insurance allowance for you and dependents • Company-provided equipment • Wellbeing, always-be-learning and home office allowances • Frequent team get togethers • Great diverse & inclusive people-first culture

Apply Now

Similar Jobs

🕒 January 28

Helix Workforce

11 - 50

💼 Consulting

📦 Logistics

🎯 Recruiter

Junior/Mid-level CRM Developer responsible for designing and maintaining CRM software solutions. Join a dynamic team to develop applications for Android and iOS platforms.

🕒 January 14

VirtusLab

201 - 500

💼 Consulting

🏭 Manufacturing

📦 Logistics

Business Relationship Developer focusing on IT tools and software development life cycle within a tech-driven sales team. Seeking candidates with AI and Security experience in a friendly atmosphere.

🗣️🇵🇱 Polish Required

🕒 January 5

Sailor Health

11 - 50

⚕️ Healthcare Insurance

🏥 Healthcare

📡 Telecommunications

Founding Engineer developing AI systems for senior mental health solutions. Leading technical decisions and engineering culture at a rapidly growing startup.

🕒 December 24, 2025

HyTechPro

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior Developer building Microsoft D365 CRM solutions, collaborating with project managers and functional consultants. Supporting high-profile clients in various industries, including sports organizations.

🕒 December 20, 2025

A1FED

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

IBM System Programmer responsible for installing, maintaining, and troubleshooting z/OS systems. Seeking a proactive individual with strong problem-solving skills and mainframe experience.