An anonymous learning platform in production for Johns Hopkins
Summary
A dedicated learning platform for Help Wanted México, the child sexual abuse prevention program of the Moore Center at the Johns Hopkins Bloomberg School of Public Health, adapted for Mexico. Nobody can know who took the course, and the client still gets the completion analytics it has to report. Zero personal data of learners, verified by audit; twelve days from an empty repository to production; operated for cents a month.
Context and constraints
A five-session online course for adults. Given the subject, learner anonymity is a program requirement: no registration, no email, no third-party cookies, no IP address in any log. At the same time the client must know how many people start, how many finish and where they drop off. The content is authored by the client in its own tool, remains the client's property and has to run untouched. Under a twelve-month work order the platform is an operated service: the client sees data and reports; the code stays with the builder.
The problem, and where it started
Off-the-shelf learning systems assume an account. The ones that allow guests still log the IP address and the browser on every request, set a cookie, or hand the video to a third party. A privacy policy cannot fix that; the code has to. The baseline was a landing page with an embedded third-party video player and a course that had only ever run inside the authoring tool.
My role
I designed, built and operate the platform and its infrastructure, and I am the client's technical counterpart. The client's team produced and validates the course content; nobody else wrote code. Every decision below is in the project's decision log, with the alternative that was turned down.
Decisions, and what was turned down
- Anonymity in the code, not in the policy. The session identifier is the hash of a random value discarded on the spot, so not even the system can reconstruct it. It travels in a header, never in the URL, because URLs end up in server logs. Validation schemas reject any unexpected field (email, IP, user agent, location) with an error and no write. Rate limits are per session, never per IP. Discarded: a "minimal" account with only an email.
- The privacy threshold as a data type. Every count has three states: zero, hidden below N, visible at N or more. Rates exist only when numerator and denominator both reach N, because a rate next to a visible count reconstructs the hidden one. The mask lives in the aggregation layer, so dashboard, CSV, PDF and the analytical warehouse inherit it. Discarded: masking in the interface, where one forgotten component brings back counts of one to four people.
- No "unique visitors", anywhere. Counting them needs an identifier, and an ephemeral one "only for counting" is still an identifier. The client receives views, not people. Discarded: lightweight fingerprinting.
- Never patch the client's package. Packages are served from the same origin so their SCORM API works as shipped; a legacy status vocabulary is translated at the edge; session time falls back to the platform's clock. Discarded: injecting code into the package to fix a resume dialog that appeared in English.
- Video and fonts never go to third parties. Discarded: the embedded player, which brings cookies, and web fonts pulled from a third party.
- No rush to production to consume the client's expiring cloud credits. Using them would have split the infrastructure across two accounts and forced a migration later. The earlier decision was superseded in writing; it cost less and there was no migration.
Result
| Measure | Value | How it was obtained |
|---|---|---|
| Personal data of learners persisted | Zero: an anonymous hash, course states, package page identifiers, server timestamps | Privacy audit of the code that writes data, real HTTP responses, real database documents and platform logs |
| Architecture decisions on record | 25, one superseded cleanly | Project decision log, delivered to the client |
| Empty repository to production | Twelve days | Architecture, functional specification and a SCORM proof of concept were closed first |
| Automated tests | Over 200 in 37 files: unit, integration against emulators, end to end including tab close | Count in the repository at go-live |
| Video egress cost | Zero, against an estimated 100 to 160 USD a month from the main cloud | Object storage with free egress behind the CDN |
| Content published | Five sessions as SCORM 2004, each in an immutable versioned path | The client keeps its source files and can take the course elsewhere |
Also delivered: daily backups, uptime check with alert, security headers, a cloud budget with alerts from day one, a runbook, and administrator access by email link over an allowlist re-read on every request. Source: the project's privacy audit report and decision log; nothing from the client's data appears here.
What did not work
The data model was clean; the leaks were elsewhere. The audit found that the platform's own request logs kept IP address and user agent, and that the published content pulled fonts and files from third parties. Both were fixed and became a pre-publication checklist. After launch, the CDN injected its own analytics beacon without the code declaring it; it was switched off the next day. And the authoring tool renumbers its pages on every publish: a learner mid-session when a new version went live got a resume dialog pointing to a page that no longer existed, and the tab closed. Reproduced twice, fixed by storing the content version with the saved position and withholding the position when the versions differ. Rule kept: audit the response actually served, not just the code.
Try it
The privacy threshold that protects learners in this platform runs in your browser in the privacy threshold in action: move N and watch counts, rates and the export hide together. The data there is synthetic; the rule is the one in production. The table above and the audit it cites remain the artifact.
What is not disclosed
Nobody on the client's team is named. No contract amount, and no vendor, project or domain name of the infrastructure appears here. The public program site and the course are linked with the client's knowledge; nothing else of the infrastructure is.
How this maps to Canadian frameworks
Canada's privacy regulator treats IP addresses and tracked browsing as personal information, and has sanctioned a dataset described as "anonymized" that still retained IPs; this platform's audit tested exactly that point, in the logs and not only in the database. Quebec's Law 25 requires opt-in consent for tracking, so a design that collects nothing identifiable needs no consent banner, and this one has none. The Ontario Information and Privacy Commissioner's principles for trustworthy data and AI use (valid, safe, private, transparent, accountable) are answered here by an audit that reads real responses and by 25 decisions on record.