Session librarydataimago

Methods · Damian Betebenner

Audit sheet: two-dimension recode (6 Oct 2026)

Open the original · source at the pinned commit · 9 passages

  1. audit-two-dimensions.0.99f647dcline 1

    Audit sheet: two-dimension recode (6 Oct 2026)
  2. audit-two-dimensions.1.68bb9947lines 3–17

    Audit sheet: two-dimension recode (6 Oct 2026) 1. The 11 items where the two coders disagree (n = 98 sample) | id | title | coder 1 | coder 2 | coder 1 rationale | coder 2 rationale | |---|---|---|---|---|---| | P008 | What Are We Assessing in Human-AI Interaction? (Keynote: Dragan Gaševi | product only | neither | Assessing learning when performance is produced through human-AI interaction in tasks. | Assessing human learning amid AI use; AI not product nor professional tool | | P036 | Rethinking Pilot Data: Evaluating LLM Synthetic Data for Scale Develop | both | profession only | LLM synthetic respondents for early psychometric analysis in scale development. | LLM synthetic data for psychometric scale development analysis | | P073 | Using Generative AI to Simulate Item Responses by Skills Insight Abili | both | profession only | GPT simulating item responses by ability band for early item evaluation. | GenAI simulated responses for early item evaluation before field testing | | P138 | Introduction to AI for Measurement | both | profession only | Title only: introduction to AI for measurement. | Title only: introduction to AI for measurement | | P143 | Building Responsible AI Practice in Language Assessment: Anticipating | both | profession only | Title only: responsible AI practice in language assessment. | Title only: responsible AI practice in language assessment | | P155 | Human–AI Collaboration in Educational Measurement: Transforming Assess | both | profession only | Title only: human-AI collaboration in measurement turning assessment data into action. | Title only: human-AI collaboration in measurement work | | P233 | Measuring Collaborative Reasoning with LLMs | product only | profession only | LLMs classifying collaborative reasoning in student discussions. | LLMs labeling student discussion reasoning; coding of discourse data | | P257 | S2A3: Thompson Sampling and Stochastic Exposure Control for High-Stake | product only | neither | Bayesian Thompson sampling adaptive testing with continuous calibration. | Bayesian adaptive testing and exposure control; no clear AI | | P282 | The Analysis of Artificial Intelligence’s Decade-Long Impact on Educat | both | profession only | Content analysis of AI's decade-long evolution in educational measurement. | Synthesis of AI's impact on the measurement field | | P283 | Transformer-Aided Detection of Gaming in Constructed-Response English | product only | both | Transformer detection of gaming in automatically scored responses. | Transformer detection of gaming in constructed responses; scoring and QC | | P304 | SarphieSense: A Pipeline-Based AI Chatbot for Navigating Living System | neither | profession only | Title only: AI chatbot for systematic
  3. audit-two-dimensions.2.4187c9dblines 5–17

    review evidence; not assessment or measurement work. | Title only: chatbot navigating systematic review evidence for researchers |
  4. audit-two-dimensions.3.96f42b0dlines 19–76

    review evidence; not assessment or measurement work. | Title only: chatbot navigating systematic review evidence for researchers | 2. Old 'product' (A) items now coded profession only (coder 1) | id | title | rationale | |---|---|---| | P010 | AI-Enabled Quality Assurance for Multiple-Choice Assessment Items | AI-enabled quality assurance of multiple-choice items: AI as item reviewer. | | P011 | Predicting Item-to-RPLD Matches with Structured AI Reasoning | LLMs classifying items to performance level descriptors to support assessment design. | | P012 | Can AI Detect Item-Metadata (Mis)alignment? An Empirical Study Using State Asses | LLMs classifying items to standards for assessment quality assurance. | | P013 | Achievement-Level Descriptor Conditioned Option-Selection Probability Modeling w | LLM-supported item evaluation predicting option selection and item difficulty. | | P038 | Prior-Informed 3PL Calibration: Reducing Sample Size via Predictive Modeling | Predictive item-parameter priors to reduce calibration samples; AI implied by session. | | P051 | Putting it Together: Using Embeddings Clusters to Identify Inconsistent Raters | Title only: embedding clusters to flag inconsistent human raters; operational quality control. | | P056 | Exploring AI-driven Methods for Pre-calibrating Difficulty of Mathematics Items | AI pre-calibration of item difficulty replacing field-testing labor. | | P058 | What LLMs Capture and Miss About Cognitive Sources of L2 Item Difficulty | LLM zero-shot item difficulty estimates versus cognitive attributes. | | P059 | Math Item Difficulty Prediction with Multimodal Input | Multimodal LLMs predicting math item difficulty. | | P060 | Beyond Item Text: Modeling Latent Cognitive Demand for Math Item Parameter Model | Modeling cognitive demand to improve item parameter prediction. | | P087 | LLM-Based Pairwise Judgment for Math Item Parameter Modeling using Workflows and | LLM pairwise judgments to recover item parameter scales. | | P088 | LLM as Investigator and Judge in Pairwise Comparisons for Item Parameter Modelin | LLM investigator and judge for item parameter modeling. | | P089 | Validating LLM Rating of Item Features for Explanatory Item Response Model . | LLM rating of item features for explanatory IRT difficulty prediction. | | P107 | Partial Identification with Multiple Nonlinear Measurements of a Latent Regresso | Estimator combining disagreeing AI scores for downstream research analyses. | | P108 | Demonstrating FairGAIte: An Agentic LLM Tool for Detecting Construct-Irrelevant | Agentic LLM tool for fairness review of pilot items. | | P110 | Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty | LLM paired comparisons to estimate item difficulty without field testing. | | P111 | Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assess | ML models predicting item difficulty in principled assessment design. | | P112 | Consensus without Accuracy: Investigating LLMs Recovery of Item Difficulty Using | LLM pairwise comparisons recovering item difficulty for AI-assisted calibration. | | P113 | When One Benchmark Hides
  5. audit-two-dimensions.4.a249598flines 21–76

    Many Truths: Mixture IRT and LLM Difficulty Prediction | Validating LLM difficulty prediction against mixture IRT classes. | | P117 | A Multimodal AI Analysis of Teacher Praise in Online Tutoring | Multimodal AI coding teacher praise quality in tutoring research data. | | P118 | Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions | AI transcription of tutoring sessions to build research logs. | | P125 | AI-Assisted Rater Training Modules: Effects on Qualification and Scoring Quality | Title only: AI-assisted training modules for human raters. | | P128 | Using Natural Language Processing to Explore Alignment Between Skill Taxonomies | NLP analysis of alignment across skill taxonomies. | | P129 | Pre-Data Embedding Diagnostics for Construct Validity in Multi-Domain Instrument | Embedding diagnostics for construct boundaries in instrument development. | | P130 | Auxiliary Information for Semantic Clustering of Assessment Items | Transformer clustering of operational items for blueprint and SME review. | | P131 | Supporting Distractor Quality Review Through Interpretable Semantic and Lexical | NLP diagnostic tool supporting distractor quality review. | | P135 | Automatic Pre-Testing of Mathematics Items: predicting difficulty parameter with | ML automated pretesting to predict IRT difficulty for calibration. | | P137 | Using XGBoost to Construct a Vertically Skill Difficulty Scale | XGBoost model building a vertical skill difficulty scale. | | P145 | Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assess | Factor analysis comparing human and LLM respondents interpreted by experts. | | P149 | Predicting IRT Parameters for Passage-Based Reading Items Using NLP | NLP features predicting IRT parameters of reading items. | | P150 | Predicting Item Statistics Using LLM-Derived Difficulty Ratings and Response-Opt | LLM-derived features predicting item statistics for certification items. | | P151 | Anchored Bradley-Terry Calibration Using LLM Comparative Judgments | LLM comparative judgments for early item difficulty calibration. | | P156 | Understanding Spatial Cognitive Process Using Time-Embedded N-Grams Model with M | Title only: machine learning analysis of spatial cognitive process data. | | P165 | Hierarchical Attention Network Architecture for Automated Response Time Predicti | Deep learning predicting item response times from metadata. | | P168 | Fine-tuning DeBERTa-v3 to Automate Spatial Language Classification | Fine-tuned DeBERTa replacing human coders for spatial language classification. | | P187 | A
  6. audit-two-dimensions.5.b1aae3e1lines 21–76

    Specialized Multimodal Transformer for Classroom Discourse Classification | Multimodal transformer classifying classroom discourse data. | | P192 | Measurement-Informed Difficulty Priors for Competition Math Items | LLM features for item difficulty priors. | | P197 | Lost Without Translation? Multilingual Sentence Embeddings for Linguistic-integr | Multilingual embeddings replacing translation in scoring reliability auditing. | | P227 | AI-Assisted Evidence-Centered Design to Build a Blueprint for a Semiconductor Cr | Title only: AI-assisted ECD to build a credential blueprint. | | P234 | An exploration of using an AI-based approach to shorten an assessment | AI shortening an assessment compared with expert-shortened form. | | P244 | Does including images improve multimodal language models’ accuracy on item discr | Multimodal models predicting item discrimination. | | P245 | Image-Based Representation of Item Response Patterns for Test integrity | Deep learning anomaly detection for test integrity. | | P246 | Listening for Meaning: Evaluating an ASR-LLM Pipeline for Thematic Audio Analysi | ASR-LLM pipeline coding student audio themes to inform assessment development. | | P265 | Beyond Latent Factors: Machine Learning Approaches to Response Clustering | ML clustering of survey responses as alternative to factor analysis. | | P277 | Why AI Demands More Principled Assessment Design, Not Less | Title only: principled design practice for AI-assisted assessment design. | | P279 | AI-Generated Distractor Plausibility Ratings as Predictors of Mathematics Item D | Title only: AI distractor plausibility ratings predicting item difficulty. | | P280 | Predicting Item Difficulty: How PAD Features Shape Model Performance | Title only: PAD features in item difficulty prediction models. | | P295 | Design Choices in Encoder-Based IRT-3PL Parameter Prediction for Passage-Based M | Encoder-based IRT parameter prediction for item calibration workflows. | | P300 | Automated Item Evaluation: Predicting Item Acceptance and Rejection using LLM-Ge | Transformer models predicting operational item acceptance from critiques. | | P303 | EquiFrame: An AI-based Program that Assesses Disability Language in Educational | Title only: AI tool reviewing disability language in measurement instruments. | | P309 | From Validation to Resilience: AI-assisted Enemy Item Identification in Operatio | Title only: AI-assisted enemy item identification in operations. | | P311 | Predicting Race and Ethnicity DIF Using Generative AI in Large-Scale Assessment | Generative AI DIF screening supporting item review in test development. | | P315 | Fairness Sentinel: An AI Agent for Early DIF Risk Screening | AI agent screening draft items for DIF risk before field testing. | | P324 | Offloading the White Elephant: Assembling with AI-estimated IRT Parameters
  7. audit-two-dimensions.6.86ed639blines 21–87

    Inste | AI-estimated IRT parameters replacing pretesting in form assembly. | 3. The 6 norms disagreements (n = 98 sample) | id | title | coder 1 | coder 2 | coder 1 rationale | coder 2 rationale | |---|---|---|---|---|---| | P154 | To Err is not only Human: Human/LLM Collaboration in Assessment | 1 | 0 | Panel on human/LLM collaboration and error in assessment work; implies division of verification responsibilities. | Human/LLM collaboration panel; no abstract | | P173 | AI-Augmented Form Assembly from a Secure Item Bank: A Faculty Pilot | 1 | 0 | Faculty test builders using AI; examines their judgment boundaries in AI-assisted form assembly. | Faculty AI-assisted form assembly pilot; judgment boundaries tangential | | P177 | Building Teacher Capacity to Generate and Evaluate Math Items with Sma | 1 | 0 | Professional development model building teacher capacity to generate and evaluate items with LLMs. | Teacher PD for AI item generation; teachers not measurement professionals | | P277 | Why AI Demands More Principled Assessment Design, Not Less | 0 | 1 | Argues AI requires more principled assessment design; design methodology rather than professional norms. | Argues AI demands more principled assessment design practice | | P282 | The Analysis of Artificial Intelligence’s Decade-Long Impact on Educat | 0 | 1 | Bibliometric synthesis of AI research trends in measurement; field content, not norms or roles. | Charts AI's decade-long impact on the measurement field; responsible adoption | | P335 | PsyMAS: A Human-in-the-Loop Multi-Agent System for Auditable Test Secu | 0 | 1 | Human-in-the-loop test security system keeping misconduct decisions human; system design foremost. | Analyst adjudication, auditability; AI must not make misconduct decisions |
  8. audit-two-dimensions.7.bf6a158elines 89–114

    4. The 22 items coded norms = 1 (coder 1) | id | title | rationale | |---|---|---| | P004 | Integrating Generative AI into R Workflows: From APIs to Shiny Apps | Frames a training gap for measurement professionals expected to adopt AI while upholding standards. | | P007 | Writing an AI-Native Dissertation | Reimagines doctoral training: building an AI-native dissertation, a change in how researchers are formed. | | P033 | Industry Leader Perspectives: How AI is Changing K12 Measurement Roles | Panel explicitly about how AI is changing K12 measurement roles. | | P046 | Responsible use of generative AI when creating reading comprehension questions: | Argues responsible-use practice: AI item evaluations should document prompts and generation steps. | | P085 | Our Future with AI: Graduate Student Perspectives on a Changing Profession | Panel on how AI changes the profession, graduate training and career preparation. | | P086 | A Jurisdictional Scan of Automated Scoring Practices in K-12 Summative Assessmen | Scan of automated-scoring practices across state programs; describes operational practice landscape of the field. | | P109 | Human-Centered AI Applications: Bridging Learning Science, Assessment, and Ethic | Panel on governance choices sharing responsibility and authority between humans and AI in assessment. | | P125 | AI-Assisted Rater Training Modules: Effects on Qualification and Scoring Quality | AI-assisted rater training and qualification; subject is training of the scoring workforce. | | P139 | Human agency before and after Generative Artificial Intelligence in education, a | Human agency in measurement before and after GenAI; plausibly addresses professionals' authority and role. | | P143 | Building Responsible AI Practice in Language Assessment: Anticipating What’s Nex | Panel on building responsible AI practice in language assessment; professional norms. | | P154 | To Err is not only Human: Human/LLM Collaboration in Assessment | Panel on human/LLM collaboration and error in assessment work; implies division of verification responsibilities. | | P173 | AI-Augmented Form Assembly from a Secure Item Bank: A Faculty Pilot | Faculty test builders using AI; examines their judgment boundaries in AI-assisted form assembly. | | P177 | Building Teacher Capacity to Generate and Evaluate Math Items with Small LLMs | Professional development model building teacher capacity to generate and evaluate items with LLMs. | | P189 | Interweaving Assessment: A Field-Generated R&D Agenda for AI and Measurement |
  9. audit-two-dimensions.8.7708449alines 91–114

    Field-generated R&D agenda for AI and measurement; the field setting its own direction. | | P221 | Implementing Generative AI Assessment Guidelines in a Mexican University | University policy process for GenAI assessment guidelines, transparency, and faculty development; governance of AI in assessment. | | P225 | Educational Measurement as an AI-Native Profession | Panel on professional norms, training, accountability and standards of an AI-native measurement profession. | | P238 | Trust Calibration in AI-Assisted Content Verification | How SMEs calibrate trust in AI verification concerns; professional review responsibilities. | | P252 | Is AI a Foundational Competency or Does AI Elevate Aspects of Competencies? | Whether AI is a foundational competency for measurement professionals (NCME competencies). | | P253 | Don’t Throw the BAIby Out with the Bathwater: Connecting FCEM to AI | Connecting NCME foundational competencies to AI; professional competencies. | | P254 | AI Competencies in Education and Credentialing: Aligning Expectations with Evolv | AI competencies for measurement and credentialing professionals aligned with evolving practice. | | P255 | Technical AI Competencies as Measurement Competencies | Technical AI competencies as measurement competencies; professional competency framework. | | P310 | Responsible AI in Digital Assessment: Case Studies with the Duolingo English Tes | Responsible AI case studies at a testing program; governance of AI use in assessment practice. |