Effectiveness study | Mixed methods | Comparative evaluation

Does concept mapping earn its place?

A four-month controlled study of whether a visual note-taking tool helps an underserved, high-churn segment keep pace — measured on performance, engagement, and the quality of what students produced.

51

Participants compared

d = .58

Higher course interest

r = .65

Map quality → grade

98%

Coder agreement

Concept map summarizing six interview themes, with perceived advantages the largest at 76%

Project overview

International students are one of the largest and fastest-growing segments in U.S. higher education — and one of the highest-churn. Many struggle and leave before finishing. The reasons look familiar to anyone who studies an underserved user group: the experience wasn't built for them. Material moves fast in a second language, tests reward a style of test-taking they were never trained for, and the default tool — bullet-point notes — doesn't help them connect or hold onto ideas.

I wanted to know whether a different tool would move the needle. Concept mapping is a visual note-taking method: instead of a linear list, you lay out how ideas connect. The real question wasn't whether it sounds good — it's whether it measurably helps the people who need it, in a live course, over four months.

So I designed and ran a controlled comparison in an introductory economics course. One section of international students learned and used the tool; a comparable section was the baseline. The honest result is a nuanced one. The tool wasn't a cure — but the students who used it kept pace with their domestic peers on exams and participation, with no significant gap, while rating the course significantly more interesting. And the better a student's maps, the higher their final grade.

Role: Sole researcher — design, analysis, and reporting

Methods: Controlled comparison, rubric scoring, motivation survey, interviews

Measures: Quizzes, midterms, concept-map scores, participation, 15 interviews

Context: Introductory economics, fully online, four months

Foundation: My University of San Francisco Ed.D. dissertation

Where international students get stuck

Three everyday frictions in a second-language classroom — and the outcome they add up to.

01 Gone before it clicks Dense material,in a second language. 02 ? The question stays in Speaking up feelsimpolite or risky. 03 Pieces that don't link Facts recorded,never connected. 04 So they fall behind Many leave beforefinishing the degree.

My contributions

I ran this study end to end as the sole researcher — I framed the question, chose the design, built the instruments, ran the analysis, and wrote it up. Every methodological call was mine to make and defend.

I built a comparison that could actually answer the question: one section of international students learning the tool, one comparable section as the baseline, same instructor and same tests, so the tool was the only thing that differed. Then I shipped the intervention into a live course — training the instructor, onboarding students, and building six mapping assignments into the normal homework rather than bolting on extra work.

To measure the maps fairly, I adapted a four-part scoring rubric and double-rated 282 maps with a second rater, calibrating first so our scores held together. I ran the quantitative analysis myself — performance trends over time, group comparisons with effect sizes, and the correlation between map quality and final grade.

Numbers alone don't explain themselves, so I paired them with 15 interviews, coded the transcripts, and checked the coding with a second coder, member checks, and peer debriefing — so the story of why the numbers came out the way they did is as defensible as the numbers themselves.

How I measured it

A controlled comparison over four months, built so the findings would hold up under scrutiny — not just look good.

DESIGN

Controlled comparison

One section used the tool; a comparable section was the baseline. Same instructor, same course, same tests — the tool was the only difference. 26 international students vs 25 mostly-domestic peers.

MEASURES

Performance, engagement, behavior

Six quizzes and two midterms; a 15-item motivation survey (usefulness, success, interest) in English and Mandarin; six rubric-scored mapping assignments; and 15 interviews to explain the numbers.

RELIABILITY

Built in, not assumed

Two raters calibrated, then double-scored 282 maps — ICC .68, quadratic-weighted κ .68. Motivation subscales α .89–.92. Interview coding cross-checked by a second coder at 98%.

What students told me

I coded the 15 interviews into six themes. One dominated: the advantages of the tool made up 76% of everything students talked about, and 14 of 15 said they would keep using it.

Six themes from student interviews A concept map. A central node branches to six themes from 15 interviews: prior knowledge and use (6%), resources for maps (5%), perceived advantages (76%), reasons for ambivalence (6%), other note-taking used (4%), and intent to keep using (4%). Each theme has sub-nodes. Perceived advantages dominates and is highlighted in orange. 6 themes from15 interviews Prior knowledge & use6% Resources for maps5% Perceived advantages76% Reasons for ambivalence6% Other note-taking used4% Intent to keep using4% Used maps across subjects Textbook, notes & lectures Usefulness 67% Success 19% Interest 14% Map shows limited info It was a course requirement Time-consuming Annotation Bullet-point notes Highlighting text Illustration 14 of 15 will keep using it
Six themes from 15 interviews. Percentages = share of coded responses; perceived advantages dominated at 76%.

Results in detail

The story above is in the language a design team uses. The statistics below are for a reader who wants to see the tests behind it.

What I tested Result
Quiz trend (repeated-measures ANOVA)Significant time-by-group interaction, F(1,28) = 4.56, p < .05, η²ₚ = .14 — tool group led throughout
Midterms (between groups)No significant difference (Midterm 1: t = 0.66, d = 0.19; Midterm 2: t = 1.02, d = 0.28)
Classroom participationNo significant difference, t(44) = −1.74, p > .05, d = −0.50
Course interestSignificantly higher, t(49) = 1.99, p < .05, d = 0.58
Usefulness / successNo significant difference (d = 0.16 / 0.12)
Map quality → final grader = .65, p < .05 (~42% of variance)
Map quality over timeNo significant change (held in the “good” tier)

Outcomes & impact

On par

Exams & participation

d = .58

Higher course interest

r = .65

Map quality → final grade

282

Maps double-scored

98%

Interview-coder agreement

14 / 15

Would keep using it

Reflection

What worked

The controlled comparison and the interviews together. The numbers showed the tool group kept pace and engaged more; the interviews explained why. Neither on its own would have been convincing.

What I learned

The honest finding was the useful one. “Kept pace” isn't a flashy headline, but for a group that English-heavy testing usually disadvantages, closing the gap is the result that survives scrutiny.

What I'd improve

The sample was small and all from one country. I'd run it larger and across more backgrounds, and treat prior knowledge and English proficiency as factors to analyze rather than noise to absorb.