Build your online resume. Claim your username
Sourcegraph logo

Staff ML and AI Agent Systems Engineer at Sourcegraph

Remote 🌍 Work from Anywhere Full time Lead USD88,000 - USD176,000 Posted  Apply before Nov 09, 2026

Job Description

Who We Are

Sourcegraph is dedicated to bringing clarity and control to the world's most complex codebases. With the rapid acceleration of code creation through AI, the infrastructure for understanding, overseeing, and evolving this code has not kept pace. Sourcegraph provides engineering organizations with comprehensive visibility across their systems, precise context for their AI agents, and the capability to execute coordinated code changes at scale. As agentic development becomes the leading engineering paradigm, Sourcegraph offers the crucial context layer teams need to manage their codebase effectively. Our products, including Code Search, Deep Search, MCP, and Agentic Batch Changes, empower engineering teams and their AI tools with cross-repository context to confidently navigate massive codebases and implement changes across hundreds of repositories simultaneously. Leading companies like Stripe, Reddit, and Leidos trust Sourcegraph to enhance shipping speed and quality. We are backed by prominent investors like a16z, Sequoia, and Redpoint, and pride ourselves on being a globally distributed team that champions high agency, direct communication, and customer focus. Join us to build the foundational infrastructure that enables every engineering team and every agent they deploy to operate with confidence.

Hours and Location

We hire almost anywhere in the world, though we prefer candidates to reside in Europe or North America for this role. However, all qualified applicants are encouraged to apply regardless of location. Regardless of your location, working hours must overlap with EST for at least 20 hours per week.

Why This Job is Exciting

Sourcegraph is at the forefront of developing AI tools to solve major challenges in the software industry, challenges that intensify with growing codebases and increasing reliance on AI agents. The Code Understanding team is responsible for the user-facing intelligence surfaces, including:

  • Deep Search: Our agentic, multi-step answer engine that operates across an enterprise's entire collection of codebases.
  • Query Assist: Transforms natural language queries into Sourcegraph's query syntax.
  • Smart Hovers: Provides concise summaries of symbols exactly where developers need them.
  • Guided diff review and the APIs utilized by both human developers and AI agents daily.

The core of these surfaces is agent engineering, a blend of software engineering, machine learning, and statistics. Agent engineering dictates which models to use and when, how to retrieve and package context, how to measure answer quality, where to fine-tune or distill smaller models for cost and latency reduction, and how to expand a single large language model (LLM) call into a robust, multi-step agent.

As a Staff ML and Agent Engineer on the Code Understanding team, you will be the technical owner, guiding the team's direction for models, evaluations, and agentic systems. Your work will measurably improve our products by making them better, faster, and more cost-effective, while also enhancing the team's proficiency in building with models.

This is a staff-level role, requiring a technical leader who is a strong individual contributor. The primary need is an expert in production machine learning and evaluation who also constructs production agent systems. You will tackle the most challenging and ambiguous problems in this domain, establish standards for others to follow, and influence strategic direction beyond your immediate team. You will have the exciting opportunity to drive the vision for delivering the best code understanding experience on the market by integrating our deterministic, large-scale systems with AI to create unprecedented user experiences.

Key Responsibilities:

  • Agentic Systems: Design and fortify the multi-step, tool-using agent loops powering current and new agentic experiences, transforming research and experiments into reliable, observable, and affordable enterprise-scale products.
  • Pragmatic Use of Evaluations: Apply sound judgment to determine where evaluations provide value, when to use targeted smoke tests and metrics, and how to avoid misleading rigor, enabling rapid, confident progress.
  • Models: Selection, Upgrading, and Training: Decide which models to deploy, drive upgrades, and fine-tune proprietary models when beneficial.
  • Retrieval and Context Engineering: Advance how models are grounded in customer code - focusing on retrieval, ranking, context windows, and citations - to enhance answer accuracy and verifiability.
  • Cost and Latency: Treat cost and latency as essential product features. You will profile, distill, cache, and right-size models to ensure ambitious features are shipped sustainably.

You will achieve this within a small, senior-leaning team that ships rapidly, owns significant product surface areas, and benefits from streamlined product management. Engineers here directly engage with customers, frame problems, and manage projects end-to-end. You will have substantial agency over technical direction and a direct line to the impact of your work.

Within One Month, You Will:

  • Get the Code Understanding products and their model/agent pipelines operational locally, and implement your initial improvements to a model, prompt, retrieval path, or evaluation.
  • Develop a clear understanding of key AI engineering challenges and the product surfaces most affected by them.
  • Become familiar with the team and our customers, beginning to form your own perspectives on the future direction of our agentic products.
  • Join the team's on-call support rotation.

Within Three Months, You Will:

  • Take ownership of a significant agentic segment of the product, from problem framing through rollout and measurement.
  • Establish the team's process for responsibly shipping model and prompt changes, including the evaluations, dashboards, and guardrails that make quality and cost regressions apparent before affecting customers.
  • Start elevating teammates' skills in building with models through pairing, code reviews, and modeling effective agent engineering practices.

Within Six Months, You Will:

  • Be recognized as the technical authority for agent engineering and agentic systems within Code Understanding, with teammates and the broader department deferring to your judgment on model, evaluation, and agent-design decisions.
  • Have demonstrably improved the products through better answer quality, reduced cost/latency, or new agentic capabilities previously deemed unfeasible.
  • Influence the direction of the team's roadmap concerning agents, presenting evidence-backed convictions on strategic bets and enabling other engineers to execute on them.

About You

You are a staff engineer and technical leader possessing hard-earned expertise across production machine learning, evaluation, and agent systems. This high-leverage role demands your ability to make sound model and evaluation decisions for a fast-moving product, build robust production systems around them, steer technical direction, and act as a force multiplier for a talented, product-minded team. You are equally adept at reasoning about an evaluation harness, a fine-tuning run, a retrieval pipeline, and the multi-step agent loop that connects them, and you actively enhance the capabilities of those around you. You operate at a staff scope, tackling the most ambiguous, high-risk problems in your domain, delving into any required codebase, establishing standards and patterns for adoption, and seamlessly translating between engineering goals and business objectives. You influence direction beyond your immediate team, leading through technical excellence and mentorship.

Required Qualifications:

  • Owned a Production Model Lifecycle: You have personally trained or fine-tuned at least one model, guiding it from dataset construction through evaluation, production rollout, and monitoring. You can articulate your choices between training/fine-tuning and prompting/retrieval, how baselines and metrics were selected, insights from error analysis, and how production data informed subsequent versions.
  • Fluent and Opinionated Agent Builder: You have designed multi-step agentic systems and ensured their reliability, observability, and cost-boundedness. You hold strong views on where agents excel and where deterministic code or human judgment is indispensable.
  • Strong Evaluation Judgment: You construct representative datasets, meaningful baselines, useful error taxonomies, and release criteria that accurately link offline measurements to production behavior. You understand when rigorous evaluations are warranted, when lightweight smoke tests or qualitative reviews suffice, and when a seemingly precise metric is misleading.
  • Treat Cost and Latency as Product Constraints: You make measured tradeoffs between quality, latency, and cost, employing the appropriate mix of model selection, prompting, retrieval, caching, distillation, and fine-tuning, rather than defaulting to more complex models.
  • Autonomous Operation on Ambiguous Problems: Given a rough product concept, customer feedback, and a Slack thread, you can independently develop a plan, prototype, milestones, and a clear perspective on tradeoffs. You own high-technical-risk projects end-to-end.
  • Contributes Beyond Domain: As a senior individual contributor, you delve into any part of the codebase required by a problem, identify issues beyond your immediate area, and translate between engineering goals and business objectives.
  • Elevates Teammates: You mentor others through collaborative problem-solving, providing insightful design and code reviews, and fostering agent engineering literacy across the team. Investing in your teammates' growth is an integral part of your role.
  • Customer and Product-Driven: You are comfortable engaging in customer calls and feedback threads, transforming raw signals into requirements, scopes, and milestones, and providing constructive pushback when feedback could derail the product.
  • Pragmatic, Not a Perfectionist: You prioritize shipping the smallest correct solution, favor robust over complicated approaches, and maintain a high quality bar with simplicity.

Engineering Fundamentals:

  • You are a strong software engineer capable of shipping production services.
  • You are proficient with our stack (Go on the backend, TypeScript on the frontend, GraphQL, Postgres, Docker) or possess a clear ability and eagerness to learn it.
  • You are fluent with agentic coding tools and take full ownership of every line of code submitted.
  • You are comfortable in an async-first, multi-service, fast-paced remote work environment.

Nice-to-Haves:

  • You have shipped an LLM-powered or agentic developer-facing product and can articulate your experience, including successes, challenges, and lessons learned.
  • You have fine-tuned, distilled, or trained models to achieve specific cost, latency, or quality targets in production environments.
  • Experience with retrieval, ranking, embeddings, or search relevance.
  • Direct experience collaborating with enterprise customers and translating their needs into product features.
  • Experience mentoring or up-leveling engineers, particularly in enhancing a team's agent engineering fluency.

Level

This job is an IC4. More details on our job leveling philosophy can be found in our Handbook.

Compensation

We offer above-market salaries to attract exceptional talent, allowing our team to focus on building great products without financial worry. Our compensation philosophy and pay bands are transparent and accessible to every Sourcegraph teammate, ensuring an equitable, explainable, and competitive approach. Your base salary is determined by the IC4 pay band for your location zone (1-4). These pay bands are informed by market data to ensure competitive compensation globally. During the recruitment process, we will discuss the specific range applicable to you based on job level, relevant skills, experience, qualifications, and location zone.

  • Zone 2: $176,000 USD
  • Zone 3: $132,000 USD
  • Zone 4: $88,000 USD

In addition to compensation, further benefits will be discussed.

Ready to Apply?

Take the next step in your career journey.

Apply Now

You will be redirected to the company's application page

Link verified 1 day ago

💜 Please mention that you found the job on True Work From Home, this helps us grow. Thanks!