Remote LLM Benchmark Scientist - Evaluations & Validation (Lima)

Remote LLM Benchmark Scientist - Evaluations & Validation (Lima)

02 ago
|
Anyone AI
|
Lima

02 ago

Anyone AI

Lima

Anyone AI Labs is seeking a Research Scientist to own evaluation design for frontier models. You’ll design benchmarks across reasoning, coding, agents, tool-use and multi‑modal tasks, grounded in expert-verified truth and validated against multiple models.

You’ll QC results to survive buyer-side review and push for durable, informative evals as models improve. You’ll lead a team to recruit experts across coding and STEM, translate lab goals into robust evaluation pipelines, and publish public

#J-18808-Ljbffr

📌 Remote LLM Benchmark Scientist - Evaluations & Validation (Lima)
🏢 Anyone AI
📍 Lima

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: remote llm benchmark scientist - evaluations & validation (lima) / lima

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: remote llm benchmark scientist - evaluations & validation (lima) / lima