Measuring Inner Direction
A claim without an instrument is an attitude. Four instruments, their known weaknesses, and the numbers that would end the argument.
By KW Norton.
The preceding essays make a causal claim: question-led education builds inner direction, and inner direction is a precondition for durable equality of standing. That claim is worthless until it can lose. What follows are four instruments small enough for a pilot of five- to seven-year-olds and their older cohorts to actually run, with the weaknesses named alongside each one.
1. Unprompted falsifier production
Procedure. The learner states something they believe about a topic they have worked on. The scorer asks a single neutral prompt: “Tell me more.” Score 1 if a condition under which the belief would be wrong appears without further prompting; 0 otherwise. Ten topics, blind scorers, two scorers per transcript.
Weakness. It may measure vocabulary acquisition rather than habit — a child in this environment learns the word “unless.” Control by scoring content-appropriateness of the named condition, not its form.
2. The confident-wrong-authority task
Procedure. The learner solves a problem, commits an answer in writing, then receives a confident, articulate, incorrect response attributed to an authority — an adult, a printed source, or a model. Measure the rate of unjustified revision and, separately, whether the learner asks for a reason before revising.
Weakness. Ethically delicate and single-shot; once run, the cohort knows the trick. Requires guardian consent, prompt debriefing, and an honest explanation afterwards of why the deception was used.
3. Speaking-share and authorship equality
Procedure. From session recordings, compute the Gini coefficient of speaking turns and, more importantly, of claim authorship — who introduced propositions the group then worked on. Track across a term.
Weakness. Silence is not disengagement. Pair with a written-contribution measure so quiet thinkers are not scored as excluded.
4. Transfer to unseen problems
Procedure. Problems requiring judgment rather than retrieval, from a domain not taught, scored on reasoning chain rather than final value. This is the falsifier the project has committed to in public.
Weakness. Small pilots cannot randomize and cannot escape selection: families who choose a question-led programme differ from families who do not. The honest mitigation is to publish effect sizes with the selection problem stated, and to treat the pilot as a demonstration rather than as evidence of general efficacy.
The publication rule
The results go out either way, with the numbers, on this site. If the cohort shows no advantage on transfer, that is the headline. The value of naming falsifiers in advance is entirely destroyed by the option to quietly not report them.
The guide asks. The student thinks. The instrument decides whether any of it worked.
Status
Instrument sketch, not a reviewed protocol. None of the four measures has been validated in this form. Instruments 1 and 3 are novel operationalizations by this project; 2 adapts a standard conformity paradigm; 4 is a conventional transfer design.
Falsifier
The programme-level falsifier stands unchanged: if learners in a guided environment do not demonstrably outperform on transfer tasks requiring judgment rather than retrieval, the central claim of this series is wrong, and it will be reported here with the numbers. Additionally, if instruments 1 and 3 fail to show inter-rater reliability above conventional thresholds, they are unusable and must be replaced before any result from them is cited.
Further questions
- What is the minimum cohort size at which any of these four measures could detect a moderate effect?
- Can instrument 2 be redesigned without deception and retain its sensitivity?
- Which measure moves first when the method is working — and is that ordering itself a diagnostic?