Hacker News new | ask | show | jobs
by radioactivist 23 days ago
I am somewhat skeptical of this.

First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement

Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.

Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.

Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.

If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.

6 comments

Worse, because students complained about the difficulty of the AI-graded quizzes, they switch to multiple-choice questions only, which increases engagement, but after analyzing the exam results they determine that multiple-choice questions don't seem to help and add AI-graded questions back, after which engagement drops again.

That means their experiment design is partially caused by their results instead of the other way around, which is a bad situation to be in. Their statistical analysis is completely inadequate for dealing with this.

And the change in engagement suggests that there's strong selection involved. Their attempt to use midterm scores to control for selection effects is unconvincing. Why not control for whether students used the platform more when there were only multiple-choice questions? Those are the ones who self-selected out of using the AI grader.

I agree with all the criticism, but I'd like to point out that this kind of study must be done in a way that doesn't discriminate any of the students. It might even be considered unethical to withhold a tool that would be already available just to do research, once there is at least some evidence or very strong suspicion that it might be, in fact, beneficial for at least some of the learners (i.e., you cannot forbid students to read books from the library just because you want to do research). Using engagement instead of results is a common choice of proxy to acquire more data, but it can become a vanity metric quickly.
> but I'd like to point out that this kind of study must be done in a way that doesn't discriminate any of the students.

This is definitionally impossible. Testing whether the tool benefits the students that get it is an attempt to benefit certain students at the expense of others. Get over this hangup if you want to do this sort of research.

Feel free to create your own university where you can do this sort of research in a discriminatory way. Have fun with the litigation afterwards. Also, explain to the peer reviewers that they just need to get over the hangup, too, when you try to publish your research.
"Feel free to create your own university where you have the funding for petri dishes. In the meantime, I'm going to publish without them, because I feel like this tissue culture I'm growing on my lunch is statistically significant."

Just because I am not personally entitled (by funding, by ethics, by career position) to collect the necessary data to demonstrate a hypothesis, does not make that hypothesis demonstrable with lesser data.

It's indeed difficult to do good education research, and instructors prefer to just mess around without a control group instead. But the intellectually humble thing to do would be to admit that the experiment can't differentiate between a method that turns bad students into good ones and one that merely identifies good ones through selection effects. When more conscientious researchers get around to doing a proper randomized controlled trial on an intent-to-treat basis, it invariably turns out to have been the latter.
> I agree with all the criticism...
It feels to me that the venn diagram between "students that fully engaged with the material" and "students that learned well from the material" is going to basically be a circle for any teaching method.
The tricky question at the forefront of education research is, at least in my mind, trying to thread the needle between effective techniques that students don’t like, ineffective techniques that students do like, and poorly defined techniques that administrators like. And on top of all that student self-reports aren’t actually very reliable indicators of learning progress at all!
This first sentence is the best summary of educational problems I've ever read, thank you.

I think we've all had a few teachers who seem able to teach 2-10x more effectively than normal; they all seem to do it in different ways. I read recently the Gates Foundation felt they'd found nothing statistically significant in their billion+ dollars of teaching philanthropy.

The most plausible education improvement proposal I'm aware of is individual tutoring. This truly does seem to work well, and I think it's why there's so much interest in agentic tutoring. Perhaps we need a TeachBench to get some hill climbing done by the frontier labs.

>effective techniques that students don’t like

Do students not like Mastership techniques?

Students don't like being told they're bad at something. Students don't like being told to do the thing they're bad at. They especially don't like having to constantly repeat this in a loop until they've caught up with the rest of the class, while other students get to enjoy their free time.

Also, mastery learning doesn't look all that effective on standardized tests: https://www.jstor.org/stable/1170613

How do you do Mastership in a class environment? Do they leave half the class waiting for everyone else to catch up?
Yeah you'd need to compare this method with some classic method as a control. Otherwise "engagement" is just an indirect measure of how much the students studied.
Yeah, calling this an "effect size" is just nonsense, and it is alarming that educational software can get away with such poor statistical practice. I'm hoping this was just a student project.
This is a helpful explanation - am not a researcher so I have little idea how to run an unbiased, meaningful experiment (except that it takes a lot of effort and thought to run one). Useful analysis
Education research is hard hard hard. Getting clean studies, especially at large sample sizes, is extremely difficult, leading to a lot of ambiguity in results.
Thank you for the feedback! Maybe the following info will be helpful when considering our results:

1. Quiz completion is our deliberately conservative lower bound on reading compliance, and the 0.71 figure is not a claim that those 16 students each gained that much. The estimate is from a regression carried by the per-lesson slope, fit across the whole dosage distribution, and the underlying dosage-performance relationship is essentially unchanged whether or not zero-completion students are included (R² 0.091 vs 0.096). In other words, more Phosphor use is strongly associated with better performance across the whole range of usage - not just the group who completed all content.

The numbers in Table 1 show how dosage was distributed across the course. We report that across the class, the median percentage of lessons reached on Phosphor, including both students with an account and those who never logged in, was above 90%. Among platform users in particular, it was 96%.

We'd like to emphasize that for this pilot, the platform was presented to students as an entirely optional "study aid", and our adoption rates far exceed those reported in the past for optional interventions. It will be interesting to see how things go when we attach completion to the course grade, as we're thinking of doing in the fall. Past literature from interventions in college courses predicts that this will achieve far higher levels of engagement, bringing the high-dosage effect to a large proportion of the class.

2. We explicitly note this in Limitations; it's an observational study. We were unable to do an RCT for this course since it raised an ethics consideration - neither we nor the instructors wanted to deprive students entirely of a course material that could have been helpful for them. We'd love to run a randomized trial at some point though - one way to do this is a crossover, where we offer the treatment to one of two groups, then switch it over to the other midway through the trial, so that both get even treatment. Another possibility is randomly selecting students to get access to MCQ-only vs. CRQ-enabled quizzes. That being said, this mechanism of conditioning on past performance is well-known and relatively robust for observational studies of educational interventions.

3. The platform was created independently of the instructors of the course. The instructors designed their curriculum ahead of time (as had been taught for years of past offerings of the course), lectured in a conventional style, referenced the course's official textbook (Freedman, Pisani, Purves) and suggested homework problems from the textbook only. The instructional content was authored using material that every student in the course had access to, and did not feature exam questions that students were evaluated on after-the-fact.

Phosphor was not endorsed publicly by the instructors, and was spread primarily via student word-of-mouth. In fact, one instructor in this course initially believed the project would be "a waste of time" and refused to collaborate with us for the pilot. Despite this, 97% of the students in this instructor's section used the platform!

Engagement also persisted well across the full ten-week term, and two-thirds of Review attempts involved retries spaced a day or more apart — not a pattern typically produced by novelty effects.

4. Instructors wrote exams independently with their long-running FPP-based curriculum. Even if we steelman and suppose that "the platform just got students to engage" rather than truly learn, this is refuted by our result that the MCQ-only Module 2 had similar engagement but no dosage relationship. This strongly suggests that the CRQ format was a driver of the results.

As we mention in the paper, we agree that replication, especially across contexts, is a priority. For us in particular, this means not only across other courses, but across other institutions as well. And an RCT would certainly help lock in the causal claim.

Thanks for this. No idea why you're getting downvoted on this background.

Anything promising that helps students learn and learn more effectively is obviously incredibly valuable. My concern here is that I don't read anything trying to disambiguate students who just, you know, work harder, and therefore used whatever materials they had to hand including Phosphor, from those who work at the same level and chose not to use Phosphor.

We all recall from our days in undergrad that there are students who do nothing and slide by, a few who do nothing and ace everything, and students who outwork everyone else and outperform -- I think a tool that only helps grinders who were already going to grind is likely not the contribution you're hoping to make.