Senior Research Engineer, LLM Evaluation and Behavioral Analysis
Questa offerta fa per te?
Crea il mio CV Crea il tuo CV e scopri la tua percentuale di corrispondenza con questa posizione — e con tutte le altre.
La posizione
Lead the design and iteration of evaluation frameworks that measure LLM behavior across instruction following, reasoning, tool use, and multi-turn interactions.
Develop specialized suites for function calling, including argument correctness, schema adherence, tool selection, multi-function planning, and error recovery.
Build agentic workflow evaluations that test task decomposition, multi-step planning, self-correction, and autonomous tool-use sequences.
Create CI/CD pipelines for A/B comparisons, regression detection, behavioral drift monitoring, and adversarial probing.
Curate high-quality evaluation datasets with nuanced and challenging cases across domains to guide model improvements.
Collaborate with researchers, engineers, and product teams to diagnose failures, triage regressions, and visualize behavior changes across releases.
Develop specialized suites for function calling, including argument correctness, schema adherence, tool selection, multi-function planning, and error recovery.
Build agentic workflow evaluations that test task decomposition, multi-step planning, self-correction, and autonomous tool-use sequences.
Create CI/CD pipelines for A/B comparisons, regression detection, behavioral drift monitoring, and adversarial probing.
Curate high-quality evaluation datasets with nuanced and challenging cases across domains to guide model improvements.
Collaborate with researchers, engineers, and product teams to diagnose failures, triage regressions, and visualize behavior changes across releases.
Vedi l'annuncio completo
Mansioni, profilo, competenze e vantaggi — crea il tuo account gratuito.
o
Hai già un account?
AccediOfferte simili
Altre posizioni che potrebbero interessarti.
Lavoro da remotoPartial
CittàSan Francisco, Stati Uniti