TixelJobs
O
Onboardmeetingsvia Greenhouse

Senior AI Engineer - Observability

United StatesPosted 3d ago
ML EngineerSeniorFull-time

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Senior AI Engineer - Observability at Onboardmeetings. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role

Sr AI Engineer (Observability) 

Function: Engineering 

Reports to: Director, Software Engineering 

Location: Remote - United States 

Position Summary 

OnBoard is a board intelligence platform trusted by over 6,000 organizations to simplify governance, and we are building AI-powered product experiences across RAG pipelines, semantic search, summarization, and emerging agentic workflows. We are looking for a Senior AI Engineer to help make those experiences reliable, measurable, cost-effective, and safe in production. 

This is a hands-on engineering role for someone who wants to build AI systems, not just observe them. You will partner with product engineering teams to design and improve AI features, instrument them deeply, and — critically — define how we know those features actually work. You will build the evaluation pipelines and datasets that tell us whether an AI experience is good enough to ship, analyze real-world behavior, and turn production signals into product and architecture improvements. 

You will help define how OnBoard ships AI: how we test prompts and retrieval quality, detect regressions, monitor cost and latency, evaluate user-facing quality, and safely evolve models, prompts, datasets, and providers over time. 

The right person is a strong software engineer with practical LLM application experience and a genuine quality mindset. You care not only that an AI feature works in a demo, but that it performs consistently for customers, degrades gracefully, provides traceable results, and improves through feedback loops. And you have real opinions about what makes an evaluation trustworthy — not just that one exists, but whether it measures the right thing. 

What You'll Do 

Build and improve AI-powered product systems 

- Partner with product engineers to design, build, and improve AI features using LLMs, RAG, semantic search, and agentic workflows 

- Contribute directly to production codebases in Python, C#/.NET, or related technologies 

- Improve prompt, retrieval, context assembly, ranking, grounding, and response-generation patterns 

- Help teams make practical architecture tradeoffs across quality, latency, cost, privacy, and maintainability 

- Support model and provider evaluations, migrations, fallback strategies, and rollout plans 

Make AI quality measurable — define what "good enough" means 

This is a defining pillar of the role, not an afterthought. It is not enough to have an evaluation; you will be responsible for whether our evaluations are adequate. 

- Design and implement evaluation pipelines for LLM-powered features, and build scoring methodologies from first principles rather than reaching for the nearest metric 

- Build and maintain versioned golden datasets covering real-world use cases, edge cases, failure modes, and customer-critical workflows 

- Implement LLM-as-judge, heuristic, human-feedback, and task-specific quality scoring approaches 

- Establish the criteria that determine whether an existing evaluation is sufficient for a given feature and risk profile — and identify gaps before they become production issues 

- Establish prompt and retrieval regression testing as part of the development lifecycle 

- Define quality gates and thresholds that help teams know when an AI feature is ready to ship 

Own observability, reliability, and cost signals 

- Instrument LLM interactions, RAG pipelines, tool calls, and agent workflows using observability platforms (OnBoard currently uses Arize; comparable tools include Langfuse, LangSmith, W&B, and OpenTelemetry-based stacks) 

- Track latency, token usage, cost, retrieval quality, groundedness, failure modes, safety signals, and user feedback 

- Build dashboards and alerts that surface meaningful product and engineering signals, not just raw telemetry 

- Analyze production traces to identify quality issues, cost spikes, regressions, and improvement opportunities 

- Create runbooks and response patterns for common LLM and AI-product failure modes 

Scale AI quality across teams 

- Build reusable libraries, SDKs, templates, and reference implementations that make correct AI instrumentation and evaluation easy 

- Document standards for tracing, metadata, prompt/version tracking, evaluation, cost reporting, and incident response 

- Coach product teams on AI quality, evaluation design, observability, and reliable release practices 

- Help establish shared patterns that let OnBoard scale AI development across product lines 

Support responsible and compliant AI delivery &l

Share