Senior Evaluation Expert – Large Language Models & AI Agents

SeniorOn-site
CompanyPatsnap
LocationShanghai, Shanghai Shi, China
Category-
DepartmentEast - Data & Technology
SenioritySenior
WorkplaceOn-site
Posted2026-09-20
Viarecruitee

Description

About PatSnap

PatSnap is a global enterprise SaaS company with extensive professional data, knowledge assets, and real-world use cases across patents, scientific literature, chemistry, materials science, and life sciences.

We are integrating large language models, RAG, and AI agents with professional data to support complex workflows including technology research, intellectual property analysis, scientific intelligence, and R&D decision-making.

We are looking for a Senior Evaluation Expert to build the evaluation system for our LLM- and AI agent-powered products and lead the team in providing reliable quality signals for model improvement, product iteration, and production deployment.

Responsibilities

-
Own the evaluation framework for PatSnap’s LLM- and AI agent-powered products, covering foundation model capabilities, retrieval and RAG, tool use, agent planning and execution, and end-to-end task performance.

-
Develop a deep understanding of professional workflows across patents, scientific research, chemistry, materials science, and life sciences. Work with product, AI, and domain teams to translate complex business requirements into measurable and reproducible evaluation criteria.

-
Build and continuously improve benchmark datasets, real-world user task sets, challenging and adversarial cases, and regression test suites, together with standards for annotation, quality assurance, and version management.

-
Combine human evaluation, rule-based methods, LLM-as-a-Judge, online experiments, and user feedback to assess accuracy, completeness, professional quality, traceability, instruction following, hallucination, safety, and task completion.

-
Establish process-level evaluation and diagnostic methods for AI agents. Analyze task understanding, planning, retrieval, tool use, and response generation to identify root causes and drive improvements across models, data, retrieval, and product workflows.

-
Integrate evaluation into development, model training, release, and production monitoring workflows, including automated evaluation, regression testing, and release quality gates.

-
Lead the evaluation team by setting objectives, allocating responsibilities, developing team capabilities, and coordinating product, AI, engineering, data, and domain teams to resolve critical quality issues.

##