Founding AI Scientist
About this job
Obsessed about training data? We're building the data infrastructure layer for AI training.
The role
As our founding research engineer, you'll own data curation as a measured discipline across every stage of training, from pre-training to post-training. Working alongside the founders, you'll turn what works into product features — and as one of the first researchers, the methods and data infrastructure you build become the foundation the company runs on.
Examples of work you might do
- Measure what's in a dataset — quality, provenance, safety, and coverage metrics that predict downstream performance
- Diagnose where data will hurt a model — decontamination, difficulty annotation, multilingual asymmetries, long-tail gaps
- Design data curation interventions — pruning, filtering, synthetic augmentation, relabeling — and attribute the gains to specific changes
- Turn the latest research into running experiments — training and evaluation loops that prove a data hypothesis
- Ship the methods that work as product features, and share findings as technical reports or papers
What we’re looking for
- Data obsession — you want to know exactly which properties of a dataset move a model, and you won't trust a result you can't measure
- Bias toward iteration — you design the smallest experiment that answers the question and move fast on the answer
- Research that ships — you optimize for models that measurably improve, not for leaderboard wins
- Product engineer — you own problems end to end and turn research into features that ship, not just results in a notebook
- Drawn to data-centric AI — you're excited by data-centric methods and the move toward autonomous AI research
Required qualifications
- Strong machine learning and deep-learning fundamentals
- Enough software engineering and PyTorch / Jax experience (or willingness to learn) to run ML experiments and build production prototypes
- Hands-on experience in one or more stages of training and evaluating LLMs/vLLMs
- Industry or research experience in one or more of: data curation, data pruning & selection, synthetic data generation, curriculum learning, dataset distillation, large-scale language/multimodal training
- Comfortable reading ML research — sourcing, vetting, and implementing promising ideas from the literature
- Able to drive applied research independently in a fast, ambiguous, early-stage environment
Nice to have
- Post-training experience — SFT, preference optimization (DPO, RLVR), or reward modeling
- Multilingual or multimodal data work
- Multi-GPU / distributed training experience
- Open-source, HuggingFace contributions
- Public technical writing, published research papers
Market insight
50% above medianExplore related jobs
Similar jobs
See all similar jobssite supervisor
WALL/CEILING PAPERHANGER
Project Engineer
hairstyle
Salon Branch Manager
Customer Service Engineer (Service Technician)
Frequently asked questions
What salary can I expect?
The employer lists 7 000 – 10 000 $ for this role at PROVENA PTE. LTD. in Singapore. For comparison, the local market median is about 5 657 $ based on 26 844 similar offers.
How do I apply for this job?
Open the original source page and contact the employer there. Finder never charges job seekers.
Are these jobs up to date?
Yes. Finder regularly refreshes vacancies from public sources and removes closed offers.
Where can I see employment type and work format?
Key conditions are shown above the description. You can also open related listings for PROVENA PTE. LTD. and Singapore.
