Educators Weigh Promise, Risks of AI Scoring Tools for Writing by Multilingual Learners
July 28, 2026 | By Katie Grant, Office of Research & Scholarship Communications
WIDA's assessments are used to identify students who need English language support ant to monitor their progress.
A new study led by UW–Madison education researchers offers one of the first detailed assessments of how K–12 teachers view the potential use of automated scoring and feedback of writing assignments by young multilingual learners.
The study, published in Assessing Writing, “Educator perspectives on automated writing scoring and feedback for young language learners: Applying a fairness and justice lens,” reports mixed reactions from teachers regarding the benefits, risks and fairness implications of AI‑powered writing tools. Even as educators expressed trust in automation’s consistency and recognized that it could be faster than human scoring, they questioned its ability to appreciate the nuances of writing by multilingual learners, among other perceived issues.
The study was led by researchers at WIDA, a research and research services organization for multilingual learners in the UW–Madison School of Education. WIDA supports state education agencies across the country with English language development standards and English language proficiency assessments of multilingual learners.
WIDA’s assessments — including Screener, ACCESS and MODEL — are used to identify students who need English language support and to monitor their annual progress in English language development. More than two million students take WIDA assessments annually.
As AI‑based writing tools become more common, the study team sought to understand educators’ perspectives before any automated scoring or feedback system is adopted by WIDA. By centering educator voices, the study provides a roadmap for designing an AI-supported writing assessment that is ethical, equitable and responsive to multilingual learners, paper authors said.
Key study highlights include:
- Concerns about bias surfaced frequently, especially related to linguistic background, digital literacy, and disability access.
- Teachers see strong potential in automated feedback, especially for timely, comprehensive responses that support learning.
- Human involvement remains essential — educators do not want automation to replace teacher judgment or interaction.
- Fairness and justice must guide system design, including transparency, piloting and alignment with standards.
The study applies language assessment specialist Antony Kunnan’s fairness and justice framework, which emphasizes equitable treatment, bias sensitivity, accessibility and positive social impact. Rather than validating a specific automated tool, researchers used educator perspectives to identify what an ethical, meaningful system would need to demonstrate.
This bottom‑up approach to validation is intentional: many K–12 teachers have limited experience with automated scoring, so their early perceptions reveal foundational beliefs and concerns that technical experts might otherwise overlook. These include trust, access, digital literacy and ecological validity — whether a system fits the realities of classroom practice.
Researchers used an exploratory mixed‑methods design, made up of:
- Phase 1: Focus groups with 16 educators across four contexts (ACCESS, Screener, MODEL, and special education).
- Phase 2: A large‑scale survey completed by 738 educators across 32 states.
Focus groups explored educators’ hopes, concerns and expectations for AI-powered writing tools. These insights shaped the survey’s design, which measured perceptions of automated scoring and feedback, bias, accessibility and washback — how assessments influence teaching and learning.
Educators raised concerns about bias related to linguistic background, writing styles influenced by students’ home languages, physical disabilities, digital literacy and typing skills. Special education teachers noted that some students who experience fine-motor challenges may experience difficulties working with a keyboard.
Educators saw promise in automation’s ability to provide quick feedback while helping to track student growth and support standards-aligned instruction. But they worried about losing the “personal touch” of human feedback, according to the paper, and emphasized that automation must enhance, not diminish, teacher–student interaction.
Teachers also offered clear guidance for how WIDA should approach any automated system:
- Pilot extensively before launching, especially for high‑stakes assessments.
- Collaborate with educators and stakeholders throughout the development process.
- Communicate transparently to build trust and understanding.
- Align automation with WIDA standards and classroom learning materials.
- Offer flexible options, including choices between human and automated scoring.
Limitations of the study included interview questions not tailored to students in specific grades and uneven participant numbers across the focus groups. Study authors said future research is needed to investigate and transparently report supporting evidence for the study’s findings.
Paper authors are WIDA researchers Mark Chapman, Lynn Shafer Willner, Jason A. Kemp, and Ahyoung Alicia Kim, plus Jieun Kim (a former WIDA intern) from the University of Hawai‘i at Mānoa (UHM).
About WIDA
WIDA provides language development resources to those who support the academic success of multilingual learners. The organization offers a comprehensive, research-based system of language standards, assessments, professional learning and educator support. WIDA’s comprehensive system is used by members of the WIDA Consortium, a U.S.-based collaborative group of 42 member states, territories and federal agencies. Learn more at wida.wisc.edu.


