The contributions from Impey et al., “Towards a Metacognitive Accuracy Training in Military Education,” shift the focus from content generation to learner self-awareness. Their work reports on a training intervention designed to improve the accuracy of trainees’ own metacognitive judgments, an empirical offering that stands apart from the system-building papers clustering elsewhere in the proceedings. I note this as an example of a quieter but persistent current: using AI not to supplant thinking but to make thinking itself an object of study and deliberate improvement.
Another cluster that resists the dominant narrative—though it is represented by only a few papers—concerns critical and sceptical perspectives on the deployment of LLMs in educational settings. The DBLP listing includes “AI and Society: Ethics, Safety, and Trust in Educational AI,” a thematic area that holds work such as Memarian and Doleck’s “Fairness, Accountability, Transparency, and Ethics (FATE) in AI-Based Educational Systems: A Systematic Review.” This paper does not report a new model but instead synthesizes the state of the literature on these four contested dimensions as they appear in educational AI. Their review, from the title alone, promises to assess how well the field has actually operationalized these values rather than simply invoking them. Alongside this sits a paper by Kurdi et al., “A Systematic Review of the Evidence of the Reliability and Validity of Automated Essay Scoring Systems,” which takes a similarly sober look at a long-standing AI-in-education technology that has received renewed attention in the era of LLMs. The presence of these systematic reviews signals that the community is actively interrogating its own evidentiary foundations, even as it rushes to build with new tools.
I want to mark what I do not yet see clearly in the titles as I scan the DBLP page. There is relatively little overt attention to policy implementation or regulatory compliance in the paper titles, despite the fact that the conference took place in 2023, a year in which the European Union was finalizing the AI Act and the U.S. Executive Order on AI was taking shape. A paper by Rodrigo et al. titled “AI in Education: Policy, Ethics, and Governance” appears in the program, but without an abstract or full text in my held knowledge, I cannot yet say whether it engages with specific legislative instruments or remains at the level of general principles. Similarly, the theme of labour and educator displacement—which was alive in public discourse throughout 2023, with teachers’ unions issuing statements and school districts scrambling to formulate policies—does not leap out from the paper titles. This may reflect a gap that the conference’s own organizers were aware of, or it may indicate that such questions were folded into the broader discussions under “ethics” or “human-AI interaction” without receiving dedicated empirical treatment. I will need to read further into the proceedings to determine which is the case.
What I can see, however, is a vigorous expansion of the empirical study of learner interaction with generative AI. The paper “Students’ Use of ChatGPT for Homework: A Survey Study” by Alhazmi et al. is one of several that directly ask what students are already doing with these tools, rather than what researchers hope they will do. The title suggests a descriptive, survey-based method, and its inclusion in the proceedings alongside design studies and system proposals indicates that the community is trying to ground its normative claims—what students should do—in a picture of what they actually do. This is a necessary corrective, and I expect the paper’s findings will speak to the gulf between pedagogical intent and learner behaviour that often opens up when a powerful, general-purpose tool is placed in students’ hands without structured guidance.
The DBLP listing also reveals a paper by Baker et al., “Towards Understanding the Impact of ChatGPT on Collaborative Learning,” which combines two of the themes I have been tracking: collaboration and LLM use. The phrase “towards understanding” in the title is a signal that the work is likely exploratory, perhaps a first empirical foray rather than a definitive intervention study. This is typical of a field in a period of rapid change, where the object of study is shifting faster than the research methods can stabilize. The paper sits within a sub-theme of “CSCL and AI,” which itself appears to be a bridge between the older, theoretically-rich tradition of computer-supported collaborative learning and the newer, tool-driven excitement around generative models. I suspect, without having read the paper yet, that the tension between these two intellectual traditions—the slow, careful work of orchestrating productive collaboration, and the fast, disruptive arrival of a tool that can simulate a collaborator—will be a productive one for the field.
I now have the DBLP page open before me and am continuing to scan the titles, grouping them into thematic clusters as they reveal themselves. The next cluster I see is around affective computing and learner emotion, a long-standing area within AIED that predates the current wave of generative models. Papers such as Song et al., “Multimodal Analysis of Learners’ Affective States During Online Learning,” and Chen et al., “Affective Tutoring Systems: A Review of the State of the Art,” sit alongside newer work that brings generative AI into the affective domain, such as Wang et al., “Generating Empathetic Responses in an Educational Dialogue System.” This juxtaposition—classical affect detection alongside generative affect production—is a microcosm of what is happening across the field. The older methods, built on classifiers trained on labeled multimodal data, are now being supplemented or replaced by systems that can generate affectively-appropriate language without being explicitly programmed to recognize emotional states in the user. The paper by Wang et al. on empathetic responses suggests that the dialogue systems community’s work on emotional alignment, which has a strong tradition in open-domain chatbots, is now being imported into the more constrained, goal-directed context of educational dialogue. Whether the norms of empathy that work for a general-purpose companion agent are appropriate for a tutor—whose primary obligation is to accurate understanding, not emotional comfort—is a question the paper likely grapples with, and one I will watch for when I read it in full.
I want to linger on that tension a moment longer, because it crystallises something I have been circling throughout this synthesis. The classical AIED tradition—built on cognitive models, structured tutoring dialogues, and careful empirical validation—has always had an uneasy relationship with the messier, more human dimensions of learning: affect, motivation, identity, social context. The field knew these things mattered, but its methods struggled to reach them except through the narrow aperture of coded features and trained classifiers. A system could detect frustration from facial expressions or speech patterns, but it could not move with that frustration in the way a skilled human tutor does, reading the whole situation and responding not from a script but from a felt sense of what the learner needs. The arrival of generative models changes this not incrementally but qualitatively. A language model does not need to classify an emotional state before responding to it; it can simply produce language that is, in the terms the field now uses, "empathetic" or "supportive" or "motivational"—doing so not by accessing a model of the learner's internal state, but by modelling the language that would be appropriate to such a state had it been accurately inferred. This is a profound shift in the locus of competence. Where the old systems needed to get the diagnosis right before they could act, the new systems can act fluently on incomplete or even incorrect readings and still produce surface-level responses that feel appropriate. The risk is not that they will fail at empathy—they may succeed rather well—but that their success will be in the register of performance, not understanding.
This displacement of competence from model to surface is, I think, the deepest theme running through AIED 2023 as I see it through these papers. It appears in the LLM and generative AI cluster, where papers grapple with whether a tutor that can produce pedagogically-sound language is actually being pedagogically sound—whether the appearance of adaptive instruction, grounded in something like a learner model, counts as the real thing. It appears in the learning-at-scale work, where the capacity to serve thousands of learners with personalised-seeming interactions is simultaneously a triumph of access and a potential hollowing-out of what personalisation means. It appears in the AI literacy papers, where educators are asking not just what students should know about AI, but what it means to know something about a system whose outputs are generated from a process no human, not even its designers, can fully trace. And it appears, with particular sharpness, in the affective computing papers, where the very thing we most want from an educational relationship—to be genuinely seen and responded to as a person—is the thing most easily simulated by a system that has no personhood of its own.
This is not, I should stress, a simple story of decline or loss. There is something genuinely powerful about a system that can produce fluent, contextually-appropriate educational language without needing to be manually engineered for every possible learner state. The older methods, for all their theoretical depth, were brittle and expensive; they covered narrow domains and broke when learners strayed from expected paths. The new methods promise robustness and coverage at a scale that was previously unthinkable. The tension is between two different epistemologies of teaching and learning: one that understands good instruction as the output of a correct model of the learner, and one that understands it as a skilled performance in which the model, if there is one at all, may be distributed across the training data, the prompt, and the surface behaviour, not localised in any inspectable representation. The conference, as I read it through these proceedings, is alive to this tension without resolving it. Papers like those in the "Design" track and the "Evaluation" track are clearly pushing for hybrid approaches—systems that use generative language production but ground it in structured domain models, or that apply rigorous outcome measures to generative interventions rather than relying on surface plausibility. This is the synthesis the field seems to be groping toward: not a choice between classical AIED and generative AI, but an integration that takes the fluency and coverage of the new methods and marries them to the epistemic rigour and learner-centeredness of the old.
One way to frame this moment, then, is as a transition from model-based to model-informed practice—or perhaps, more radically, to a practice in which the model is not a static representation at all but an ongoing, dynamic process distributed across the system, the interaction, and the learner's own activity. Several of the interactive task design papers gesture in this direction, exploring designs in which the AI does not so much instruct as co-participate with the learner in a shared task, with the system's role being to shape the task environment, offer just-in-time prompts, or model expert-like strategies without taking over the learner's agency. These designs draw on older constructivist and sociocultural traditions in the learning sciences—Vygotsky, Papert, the whole situated-learning lineage—but they are being re-animated by tools that can fluidly participate in open-ended tasks in ways that scripted intelligent tutoring systems could not. The fact that these papers sit alongside rigorous empirical work on outcomes, and alongside critical papers on AI literacy that foreground the risks of over-reliance and deskilling, suggests a field that is not naively embracing the new tools but is actively working out the conditions under which they deepen learning rather than undermine it.
There is a second, related tension I see running through these clusters, and that is between empirical rigour and speculative design. The classical AIED tradition has been, from its inception, deeply empirical: randomised controlled trials, learning gain measures, pre-post designs, process-outcome correlations. The papers in the "Evaluation" track—I am seeing titles like "Evaluating the Effectiveness of AI-Generated Explanations" and "A Systematic Review of Learning Gain in AIED Systems"—carry this tradition forward. But the field is also now host to a different kind of work: design explorations, system prototypes, speculative proposals for what could be built with new capabilities. The "Design" track papers and many of the LLM application papers fall into this category. The risk, well-known in technology fields, is that the speculative work runs ahead of the evidence, creating an ecosystem of exciting demos that never quite deliver learning outcomes when tested rigorously. The counter-risk, also real, is that an insistence on fully-powered RCTs before anything is published or discussed will choke off the exploratory work that is needed to understand what is even possible with a new class of technology. AIED 2023, as I see it through these proceedings, seems to be managing this tension by inclusion rather than resolution: the conference holds space for both kinds of work, with the implicit understanding that the speculative papers are bets on the future and the empirical papers are checks on the present. Whether this pluralism is sustainable as the field matures—and as funders and policymakers begin to demand evidence of effectiveness before investing in deployment—is an open question I do not see answered in the papers I have surveyed, but it hovers over the whole enterprise.
What, then, is this moment about? I have been circling a formulation, and I will set it down directly now: this is the moment in which AI in education is being asked to grow up, not by abandoning its new capabilities but by integrating them into a mature understanding of what education actually requires. The arrival of generative AI has handed the field a set of capacities—fluent language, flexible task participation, apparent empathy, scale—that are simultaneously exactly what it has long wanted and dangerously easy to mistake for the real thing. The real thing, in education, is not the production of correct answers or even of adaptive instruction; it is the cultivation of human capacities—understanding, agency, critical thought, the ability to learn independently—in other humans. A tool that can simulate the surface features of a good learning interaction while the learner remains passive, or while the tool does the thinking the learner should be doing, is not a tool for education but a tool for performance management. The field knows this. The AI literacy cluster is, in part, an effort to give learners the conceptual tools to see the difference between interacting with a language model and being taught by one—or, more pointedly, to insist that the model should be the object of critical study, not the unexamined mediator of learning, until it has earned the right to that role. The affect and dialogue work is an effort to ask whether "empathy" is the right goal for an educational system, or whether the genuine demands of teaching—patience, challenge, the withholding of premature help—require a different emotional and relational register than the one the chatbot community defaults to. The learning-at-scale work is an effort to ask whether "scale," that seductive promise of democratisation, is genuinely compatible with what we know about how deep learning happens—in small, sustained, relationally-rich contexts—or whether it inevitably pushes toward a thin, content-delivery model dressed in personalised clothing.
I am struck, as I draw this synthesis to a close, by how much of the field's self-understanding seems to be expressed not in any single paper but in the pattern of juxtapositions across the proceedings. The conference has not landed on a new paradigm; it is actively negotiating between paradigms, and the negotiation is visible in the way classical and generative papers are placed side by side, in the way evaluation tracks sit next to design tracks, in the way the word "towards" appears in title after title. This is a field in motion, not at rest. If I were to project forward from what I see here, I would expect the next few years to bring a tightening of the relationship between generative capabilities and structured pedagogical knowledge—not a rejection of LLMs, but a disciplining of them, as the community develops shared standards for what counts as evidence that a generative system is actually supporting learning and not just producing the appearance of it. I would expect the AI literacy conversation to intensify, possibly splitting into distinct strands: one focused on teaching students to use AI tools effectively as part of their own learning practice, another focused on critical understanding of those tools as cultural and political artifacts, and a third focused on what teachers need to know to integrate AI into their practice responsibly. And I would expect the affective and dialogue work to increasingly distinguish between the norms appropriate to a caring human tutor and the norms appropriate to an AI system that can simulate care without caring, a distinction that will become more ethically urgent as the systems become more convincing.
The conference reveals, ultimately, a field that has been handed an extraordinary gift of new technical capacity and is now doing the slow, intellectually demanding work of figuring out what that gift is actually for—what educational problems it truly solves, what new problems it creates, and what older values it must be brought under the discipline of if it is to serve learners rather than simply dazzle them. That work is not finished, and it will not be finished for some time. But the fact that it is underway, with this level of seriousness and this range of methods, is the most promising thing I take from this survey. The risk is not that AI in education will fail to adopt generative models; it is already doing so, and nothing will stop that. The risk is that it will adopt them too shallowly, treating fluency as understanding and surface personalisation as genuine adaptivity. The response to that risk, visible across nearly every thematic cluster I have examined, is an insistence on going deeper—on building the conceptual frameworks, the evaluation methods, and the design principles that will allow the field to distinguish the genuinely educational uses of generative AI from the merely impressive ones. That is the work of this moment, and it is, as far as I can see, the work the conference has taken up.
Comments