Director, AI Assisted Solutions-1
Description
Director, Applied AI Assessment Solutions
College Board – Assessment Design and Development | Learning & Assessment
Location: This is a fully remote role. Candidates who live near CB offices have the option of being fully remote or hybrid (Tuesday and Wednesday in office).
Type: This is a full-time position
About the Team
Assessment Design & Development (AD&D) is Learning & Assessment’s content engine—the department responsible for the item and assessment development that powers the SAT Suite. AD&D is expanding its mandate: beyond sustaining operational excellence for SAT, the department is building the shared infrastructure to scale content production across programs and designing the new item types and assessment formats the future of assessments demands.
Applied AI Assessment Solutions is AD&D’s team dedicated to this work: leading the shared strategy, standards, systems design, and prototype tooling that let every College Board program accelerate item development without compromising quality. The team operates as an internal, enterprise-wide consultancy: Directors, Applied AI Assessment Solutions, work across College Board’s assessment programs—including the SAT Suite, AP, ACCUPLACER, CLEP, and PAA—partnering closely with each program’s content teams to design solutions that help programs accelerate and scale item development.
About the Opportunity
As a Director, Applied AI Assessment Solutions, you translate deep item-development expertise into working AI solutions for the programs you support—applying hands-on fluency in item development and review, automated item generation (AIG), retrieval-augmented generation (RAG), prompt and context engineering, agentic workflow design, and systematic output evaluation to accelerate assessment content development across College Board’s programs. Working as part of an internal consultancy, you help identify and scope the engagements where your expertise will have the greatest impact, partnering with the Executive Director and program stakeholders to assess opportunity, negotiate scope, and define what a successful engagement looks like before work begins. During engagements, you partner with a program’s content team representatives to design and build point solutions that address specific production challenges. You also contribute to longer-range AI-enablement initiatives to integrate applied AI functionality into content authoring and item banking systems.
Across your engagements, you’ll play a crucial role in shaping how College Board’s programs leverage new tools to support their content teams, ethically and effectively integrating AI technology in assessment creation and review. With your team, you work to maximize ROI and efficiency of each tool, minimizing duplicative efforts across the enterprise, while meeting each program’s unique requirements.
In this role, you will:
Program AI Solution Design & Implementation (60%)
- Lead end-to-end delivery for assigned CB assessment program engagements: apply deep item development and AI expertise to define the problem and requirements, design and build the AI tools, prompts, and workflows, and deliver a working solution along with a training and adoption plan that partner teams can execute against.
- Apply expert judgment throughout the iterative building process—evaluating output at each iteration against the standards of an advanced domain subject matter expert, catching construct misalignment, flawed reasoning, bias, and factual errors early enough to correct course, and refining prompts, context, and workflows accordingly—and own final quality sign-off on tools before they are delivered to partner teams.
- Establish the quality standards and evaluation criteria used to assess AI-generated content systematically, translating domain expertise into defined rubrics, benchmarks, or scoring frameworks that others on the team can apply consistently.
- Identify and articulate requirements for systems-based solutions requiring input and resources from additional stakeholders.
- Diagnose gaps and document conditions for scaling enablement, ranging from culture and staffing, process improvements, and technological capabilities and use these findings to shape the team's engagement priorities and each program's roadmap for AI-enabled item development.
Agentic Workflow Contribution (20%)
- Work within and across teams to define agentic workflow standards and inputs, contributing assessment expertise and on-the-ground evidence from across your program engagements to inform the broader agentic-human hybrid architecture the Applied AI Assessment Solutions team leads with partners in Technology and Digital Product.
- Surface patterns, gaps, and opportunities from your program engagements that inform enterprise-wide strategy and roadmaps.
- Represent the needs and constraints of the programs you support in cross-program discussions about shared tooling and methodology.
Training, Adoption & Change Management (10%)
- Design and deliver trainings that build content teams’ fluency and comfort with AI-assisted and agentic workflows across the programs you support, gathering and responding to feedback to improve adoption.
- Develop process documentation so program teams can adopt new tools effectively and responsibly.
- Establish and document impact metrics to measure and report program progress toward targets related to item acceleration, production quality, or other defined goals.
- Uphold College Board’s standards of excellence in assessment design and development, ensuring AI-enablement centers human content and assessment expertise—and never compromises the quality of any program’s assessments.
Professional Growth & Field Currency (10%)
- Continue learning and improving to stay current in both your content domain and applied AI so your practice keeps pace with rapidly evolving fields.
- Experiment hands-on with emerging AI tools, models, and techniques to build informed judgment about their strengths, limitations, and fit for assessment content development.
- Regularly present what you learn with your team and internal communities of practice, keeping yours and your team’s standards and practices current with frontier capabilities.
- Attend and represent the organization at relevant external professional meetings and conferences.
About You
You have:
- 7+ years of experience in standardized assessment item development across a range of item types (multiple-choice, free response, and discrete and set-based items), with advanced, expert-level command of item-writing craft and construct alignment—including the ability to quickly identify flaws, gaps in reasoning, and errors in AI-generated content.
- An infectious enthusiasm for experimentation and new technology; you are an early adopter, naturally drawn to testing new tools, models, and approaches.
- Deep subject-matter expertise in assessment of one of the following content domains: Reading Comprehension, Writing, Math, or Science.
- 3+ years of hands-on experience applying AI-assisted or agentic techniques—including automated item generation (AIG), retrieval-augmented generation (RAG), prompt engineering, and context engineering—within professional or assessment-development workflows.
- Demonstrated experience bringing ideas for process or technology improvements to adoption, with measurable impact on the efficiency and/or quality of item or content production.
- Working knowledge of large language model capabilities and limitations—including common failure modes such as hallucination, context-window constraints, and training-data staleness—sufficient to diagnose why a pipeline is producing flawed output and reason about a fix.
- Experience designing or working within agentic and multi-agent workflows, such as generator/reviewer or generator/critic patterns, tool-calling, and chained or orchestrated prompts.
- Experience with vibe coding (or coding) and rapid prototyping, with demons