anyone-ai logo

Research Scientist

anyone-ai
Remote
RemoteUruguayChileUnited KingdomEcuadorMexicoColombiaFrancePeruBrazilArgentina· 2 months ago

Straight from anyone-ai’s careers page. Apply on the company site — no recruiter, no middleman.

Explore more remote Research Scientist jobs — salaries, top companies, and the latest openings.See all →

Research Scientist (Remote/US/LATAM)

Department: Anyone AI Internal Team

Location: Fully Remote, Uruguay, Chile, United Kingdom, Ecuador - Fully Remote, Mexico - Fully Remote, Colombia - Fully Remote, France, Peru - Fully Remote, Brazil, Argentina - Fully Remote

Employment Type: Contract

Research Scientist, LLM Evaluations & Benchmarking

Anyone AI Labs
Reports to: CEO · Remote / LatAm / US

The role Evaluation is one of the hardest open problems in AI: we still dont have reliable ways to measure what frontier models can and cant do, and the field mostly runs on benchmarks that are saturated, contaminated, or measuring the wrong thing. Youll own that problem at Anyone AI, measuring frontier model capability.

This is a research role at heart: you decide what a good evaluation is , design the benchmarks that prove it, and defend the methodology under lab scrutiny. Youll build frontier-grade evaluation packages across reasoning, coding, agents, tool use, and multi-modal — grounded in expert-verified truth, validated against multiple models, and QCd to survive buyer-side review.

Responsibilities

● Evaluation research. Turn eval targets into original benchmark designs. Own the hard measurement questions: construct validity, item discrimination, headroom, reliability, contamination, and capability elicitation. Push toward evals that stay informative as models improve.

● Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).

● Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.

● Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what theyre trying to measure and translate it into an evaluation design.

● Delivery & dissemination. Turn lab requests into winning sample packages and own pilots end to end. Where the work generalizes, help turn it into public benchmarks and papers: we support publishing at venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.

What were looking for

● Research background in ML evaluation or benchmarking (a track record of published or open benchmarks, eval/measurement research, or equivalent hands-on work that labs have relied on).

● Deep LLM/frontier-model benchmarking expertise, with real strength in code-model and agentic evaluation.

● Fluency with the measurement problem itself: construct validity, psychometrics, rubrics, pass rates, headroom, contamination, and what makes a task genuinely discriminate a model.

● Interest in the safety side of evaluation, capability elicitation, robustness, and measuring the things that are hardest to measure honestly.

● Proven ability to hold a team or expert pool to a rigorous standard.

● Comfort with the full research loop: framing the question, running the study, and writing it up.

● Fluent English; Spanish a plus.

Path Robotics logo

Path Robotics

Senior Recruiter, AI/ML Research (Remote)

Remote
✓ From careers page· about 2 hours ago
Clario logo

Clario

Imaging Research Associate

Remote
✓ From careers page· about 3 hours ago
Nielsen logo

Nielsen

Market Research Analyst (Remote)

Remote
Mexico City, MX
✓ From careers page· about 7 hours ago
Alcanza Clinical Research logo

Alcanza Clinical Research

Contracts Administrator

Remote
Lake Mary, FL
✓ From careers page· about 7 hours ago

Discover More than 100,000 Hidden Remote Jobs Before Everyone Else

Unlock All Remote Jobs Today

Simple pricing. Big savings on Quarterly and Yearly.

Monthly Access

$19/month
  • Instant access to fresh remote jobs from 500+ companies
  • New opportunities added hourly, often 3-7 days before anywhere else
  • Advanced filtering by role type, stack, pay, and location
  • Priority customer support
Start 7-day trial — $2.95
Most Popular

Yearly Access

$59/year
  • Everything in Monthly
  • Save $169 (~74%) vs paying monthly
  • Average job search takes ~6 months - get covered for the whole journey
  • Less than the cost of one lunch per month for competitive advantage
  • Equivalent to just ~$4.92/month
Start 7-day trial — $2.95