> For years I've been trying to get any proctoring company to agree to a study where I try to cheat. None have agreed. I've had legal advice not to do such a study without permission from the vendors.
I'm guessing they had to fly under the radar a bit.
That's because defeating the webcam proctoring is trivial. The simplest setup is a teensy, allowing two mice to be shown to the system as one input device. A HDMI EDID emulator placed after the GPU and before the splitter for two monitors.
The second person literally doesn't have to be in the same room, and you could even get fancy with a PiKVM for your assistant to be remote.
The proctor cannot distinguish the two mice or monitors, because the computer cannot either. This is a completely unpatchable hole with no good method of detection.
The standard for defeating this is a 2-device setup on zoom or similar: one phone and one computer, both with cameras. Phone must provide some side view of the surrounding environment.
It's multiple choice questions, all you have to do is not be blatant about it. The helper only needs to add a bit of side movement as the test taker brings the mouse down the list of possible answers.
Otherwise the test taker is in control 99% of the time. The helper doesn't even need to ever click or meaningfully move the mouse.
You could go another level and do a single earbud. Maybe the check for those now that AirPods are super popular, but they sure never needed to see my whole head when we I took remote proctored exams five years ago.
I've had time to think about this, just never bothered to act because the exams I was taking were a joke anyways and I was just looking to get my degree and move on, not make a point on how easy it would be to beat the system.
Absolutely. From the article: the software detected 0/6 cheating students, while a human detected 1/6 cheating students.
Frankly, with such low numbers, I would not draw any conclusion at all. Within the margin of error, the human could have detected 0, or maybe 2 cheating students. Who knows. It would change this results dramatically, so you can basically ignore them.
Also, the cherry on the cake is that the human also detected one cheating student who wasn't cheating. So human vs. software, no one wins.
First, I did not find any statistical test in the paper. The sample size is not relevant.
Further the power of a test already required you to know two parameters: variance and the size of the effect your testing for (difference between means). Proctorio is very careful in not claiming any effectiveness for detecting cheaters, how will you estimate these parameters?
Also iirc, the rule of thumb was for a normal distribution 30 samples is enough for decent strength, 40 samples is enough for general distributions unless you're looking at very small effects, weird distributions or in very noisy experiments.
Right, I'm trying to do this sorcratically and we got a smarty-pants over here. HN constantly has sample size critiques without power analysis responses which is a shame.
There are very well documented formulas for calculation sensitivity and specificity of a diagnostic with a certain degree of certainty around it. You would use one of those.