justbacked
← All jobs
Vinyl Equity

AI Evaluation Engineer

Mangaluru Office · Full-time

Series A · $20M · 124d ago Posted 2d ago
Have your AI draft the answers and brief you on Vinyl Equity→

About the Role

As an AI Evaluation Engineer, you'll own the quality bar for both our AI-powered features and the broader product experience they sit inside. You'll design and run evaluation frameworks that catch regressions in model behavior (accuracy, hallucination, safety, compliance-sensitive edge cases) as well as classic product/QA issues — so that what we ship to customers handling real securities and compliance data is trustworthy every time. This is a hands-on, blended role: part evaluation engineering for LLM/AI systems, part product quality ownership.

What You'll Do

  • Design, build, and maintain evaluation frameworks and benchmark suites for AI/LLM-powered features (e.g., document extraction, compliance checks, automated workflows), covering accuracy, consistency, hallucination rate, and safety.

  • Define golden datasets, rubrics, and scoring methodologies (human-in-the-loop and automated/LLM-as-judge) to measure model and product quality objectively.

  • Build automated eval pipelines that run in CI/CD, flag regressions before release, and produce clear, trackable quality metrics over time.

  • Extend evaluation coverage beyond the model layer into full product/QA testing — functional, regression, and end-to-end testing of AI-powered features and the surrounding product.

  • Partner with product managers and engineers to translate ambiguous quality bars ("is this good enough to ship?") into measurable, repeatable evaluation criteria.

  • Investigate failures and edge cases, perform root-cause analysis across the model/product boundary, and drive fixes with engineering.

  • Maintain traceability and reporting on eval/QA results for compliance-sensitive workflows, given the regulated nature of the data we handle.

What We're Looking For

  • 5 plus years of experience in QA/test engineering, ML evaluation, or a related quality-focused engineering role.

  • Hands-on experience testing or evaluating AI/LLM-powered features — building eval sets, scoring rubrics, or benchmark harnesses (or strong adjacent automation/QA experience with a demonstrated interest in AI evaluation).

  • Solid automation/testing fundamentals: scripting (Python and/or TypeScript/Java), API testing, SQL for data validation, and CI/CD integration.

  • Comfort working with LLMs and AI tooling directly — prompt engineering, RAG pipelines, or AI-assisted development tools — either as a builder or a rigorous evaluator of them.

  • Strong analytical mindset: comfortable defining metrics for fuzzy, subjective quality questions and defending them with data.

  • Clear written communication — you'll be documenting failure modes and quality bars for both engineers and non-technical stakeholders.

  • Bonus: experience in fintech, compliance, or another regulated domain where correctness and auditability matter.

Nice to Have

  • Experience with vector databases / semantic search (FAISS, ChromaDB, or similar) and RAG evaluation.

  • Familiarity with eval tooling/frameworks (e.g., Ragas, DeepEval, promptfoo, custom LLM-as-judge pipelines).

  • Background in Agile/Scrum environments and defect-management tooling (Jira or similar).

What We Offer

  • Competitive compensation based on experience

  • Equity participation

  • Comprehensive health insurance (self + family)

  • Paid leave and wellness benefits

To apply: send your résumé to [email protected].