דילוג לתוכן הראשי

הגדרות עוגיות

בחרו מה לאפשר. אפשר לשנות את הבחירה בכל עת דרך „הגדרות עוגיות” בתחתית כל עמוד.

חיוניות

התחברות, אבטחה ושמירת הבחירות שלכם באתר — כולל הבחירה הזו, ערכת הצבעים והגדרות הנגישות.

תמיד פעילות

סטטיסטיקה ומדידה

אילו דפים נצפים ואיך משתמשים בהם, כדי לשפר את האתר. בלי שם ובלי פרטי קשר.

כלים: Google Analytics, Microsoft Clarity

פרסום ושיווק

מודעות שמותאמות לתחומי העניין, והתראות דחיפה על משרות חדשות למי שביקש. בלי הסכמה מוצגות מודעות כלליות בלבד.

כלים: Google AdSense, OneSignal

פרטים נוספים במדיניות הפרטיות.

JOBTIME
לוגו Team8

Team8- Fintech Stealth Startup- Senior AI Researcher, Evaluation & Agent Performance

Team8

עלתה ל-JOBTIME לפני שעה· בתוקף עד 21 בנובמבר 2026

מיקום
תל אביב - יפו, מרכז
היקף משרה
משרה מלאה
תפקיד
דאטה סיינטיסט

תיאור המשרה

Description

About us

We build the context layer for engineering organizations. We connect code, runtime telemetry, data lineage, and operational systems into one live model of how software runs, then give AI agents that model through chat and MCP. 

What you'll do

  • Own the evaluation function for all Cobalt agents: code Q&A, change-impact analysis, incident response, and MCP tool use by coding agents.
  • Design and build eval harnesses and benchmarks, with ground truth taken from merged PRs, production traces, and code.
  • Define and track agent performance metrics: recall and precision, cost per correct answer, tool-call efficiency, latency, and failure modes.
  • Run controlled comparisons across models, prompts, retrieval strategies, and baselines, with statistical rigor.
  • Diagnose agent failures in retrieval, reasoning, and tool design, and drive the fixes with engineering.
  • Build regression suites in CI that score every model, prompt, and connector change before release.
  • Lead external research, including ownership of research papers.
  • Set the research agenda and mentor future hires.

Requirements

Requirements

  • 7+ years in ML/NLP research or applied AI, including evaluating LLMs or agents.
  • A track record of designing benchmarks: task sampling, contamination control, scoring methods, and power analysis.
  • Deep knowledge of agent architectures: tool use, retrieval, multi-step planning, MCP or similar protocols.
  • Expert Python, and production-quality eval infrastructure you have shipped.
  • Experience leading research projects end to end, from design to publication or product.

Nice to have

  • Publications at NeurIPS, ICLR, ACL, ICSE or FSE on LLM evaluation, code intelligence or agents.
  • Experience with repo-level code benchmarks (SWE-bench-style).
  • Experience with LLM-as-judge methods and their known failure modes.
  • A background in distributed systems, observability data, or large codebases.

Why join

  • Enterprise ground truth: thousands of repos, live traces and production incidents.
  • Your evals decide what ships.
  • A founding role in the research function.

משרות דומות

לוגו התעשייה האווירית
חדשדאטה סיינטיסטים

אלגוריתמאי/ת

התעשייה האווירית

  • פתח תקווה
  • משרה מלאה
לוגו התעשייה האווירית
חדשדאטה סיינטיסטים

סטודנט/ית AI

התעשייה האווירית

  • אשדוד
  • משרת סטודנט
לוגו ActiveFence (Alice)
חדשדאטה סיינטיסטים

Senior AI Researcher

ActiveFence (Alice)

  • רמת גן
  • משרה מלאה
לוגו Payoneer
חדשדאטה סיינטיסטים

Senior Data Scientist

Payoneer

  • תל אביב - יפו
  • משרה מלאה
לוגו אלביט מערכות
חדשדאטה סיינטיסטים

GenAI Engineering Expert

אלביט מערכות

  • מרכז
  • משרה מלאה

הגדרות נגישות

ערכת צבעים

גודל טקסט

100%

התאמות תצוגה

הצהרת נגישות