Discussion about this post

User's avatar
QUASAR's avatar

The assessment calibration example is the part I keep coming back to.

AI doesn’t need to replace the grader to be useful. Even surfacing inconsistent scoring, recurring misconceptions, and unclear rubric language could make assessment much fairer.

The important boundary is exactly the one you’ve kept here: let the system reveal patterns, but keep the final judgment accountable to a human.

That feels like a much healthier direction for educational AI—less “let the model decide,” more “help us see what we were missing.”

Jeffrey B. Matu's avatar

This summer, working with Dr. Amine Amar, we have been testing Fikira (https://framethechangelab.com/fikira-download.html) with three faculty members in Morocco, funded by Al Akhawayn University, and because two of them teach in French and one in Arabic we have had to think from the start about a tool that cannot quietly assume everyone works in English.

Fikira is a free learning companion that runs on the teacher's own computer and keeps working when the internet drops, which matters because the teachers and students we most want to reach do not have reliable broadband or the budget for an AI subscription. A teacher gives it a syllabus, and it looks for the assignments a student could now finish with AI without doing any of the thinking the assignment was meant to develop, then suggests ways those assignments might be redesigned. We kept students out of this first phase on purpose, since it seemed wrong to put the tool in front of a class before the people who actually teach the course had told us whether it was worth their time.

The reality has been harder than the idea, mostly in ways we did not see coming. A local model has a training cutoff and has never seen the course in question, so its early advice was fluent, plausible, and could have been written about almost any syllabus anywhere. Then in June we hit something considerably more embarrassing, when the analysis timed out on a teacher's laptop and a fallback message we had written for our own debugging was printed into a generated syllabus in the course description field, where students would have read a sentence explaining that Fikira had used a fast local analysis because the full one was unavailable. Around the same time we found that the tool had lifted a sentence about grading percentages out of a syllabus and was analysing it earnestly as though it were an assignment.

What helped in the end was not the larger model we assumed we needed, but taking away the model's licence to assert anything it could not point at. Fikira now pulls the assignments out of the document before the model is involved at all, quotes the wording behind every finding it reports, and says plainly that something is not evidenced when there is nothing in the file to quote, so that when it cannot tell whether an exam is supervised or taken at home it says so rather than quietly deciding for itself.

If we could pass on one thing to anyone building something similar, it would be to test on a real teacher's syllabus rather than the tidy examples you wrote yourself, because ours passed for weeks while an actual course document was failing in ways we never saw.

We are still in the middle of this. The teacher survey went out at the end of June and we are waiting on responses, and what those teachers tell us will decide the shape of the next phase, which is a pilot with their students once the faculty are satisfied the tool is ready for them.

No posts

Ready for more?