This handbook answers the questions a researcher asks in roughly the order they become urgent: what is here, whether you need your own ethics approval, how data leaves, what you should write in your limitations section, and how to cite any of it. It is deliberately explicit about what the platform cannot give you, because finding that out in month four of a study is expensive.
1. What this platform is good for
Uedu is a production teaching platform that happens to be instrumented, not a laboratory. That determines what kinds of question it answers well.
| Good fit | Poor fit |
|---|---|
| Observational study of authentic AI-mediated learning at scale, across disciplines and institutions | Tightly controlled laboratory experiments requiring standardised conditions |
| Longitudinal work within a semester, or across semesters for returning students | Anything needing random assignment of students to institutions or teachers |
| Design-based research where you change your own course and observe what follows | Studies where the intervention must be hidden from the teacher |
| Method and instrument work: measuring dialogue, comparing analytic approaches, validating an indicator | Claims about learning gains from a single semester of a single course |
| Multimodal integration — dialogue with assessment, or with cognitive testing, or with physiological daily summaries | High-frequency physiological research needing raw continuous signals |
The framework the data is organised around
Educational Omics treats learning as multiple simultaneously observable layers. It is the organising scheme for what the platform collects, and knowing the vocabulary makes the rest of this handbook shorter.
| Dimension | What it covers |
|---|---|
| Cognomics | Cognitive process — dialogue trajectories, cognitive-level classification, cognitive test performance |
| Linguomics | Language and expression — complexity, semantics, speech transcription |
| PhysioNeuromics | Physiological and neural — heart-rate variability, sleep, stress; sensor work in development |
| Sociomics | Social interaction — forums, collaboration, peer response |
| Environomics | Learning environment — ambient conditions derived from public monitoring data |
| Ethicomics | Ethics and governance — consent, withdrawal, disclosure, and the study of those mechanisms themselves |
Prior work on the platform, including the papers that define these terms, is at /practice (19 publications). Reading two or three before the first conversation makes that conversation much shorter.
2. Which ethics case are you in
This is the first question to settle, because it determines your timeline more than anything else. The platform holds an umbrella research-ethics protocol (NTU-REC 202507EM058), and data governance under it is the platform's responsibility rather than yours. Three cases:
Covered. You do not file your own submission for this. De-identified export of data already collected runs under the umbrella protocol. Consent, de-identification and the export route are the platform's responsibility. We provide a provenance statement you can cite in your methods section (chapter 10).
Needs a protocol amendment before it starts. Plan lead time. A new instrument, a new sensor, an external questionnaire, or any question put to students that the protocol does not already cover. Amendments are batched rather than filed one at a time, so the determining factor is how early you tell us — not how urgent it is when you ask.
Yours. We supply the paperwork you need to answer them. Some institutions require their own registration or approval regardless of where data was collected, and most journals require an ethics statement. Those obligations sit with you, but they are usually satisfied by citing the umbrella protocol rather than by running a fresh review.
Where people get this wrong
- "I'll just add three questions to a survey." That is Case B. Three questions the protocol does not cover is new collection, however small.
- "It's my own course, so it's my decision." Teaching it is your decision; collecting research data from the students you grade is not the same act, and the separation in chapter 8 applies without exception.
- "I'll use an external form, it's faster." It is slower. An external questionnaire falls outside the umbrella and lands you in Case B with an extra data-transfer problem. Use the platform's own survey tool — that is precisely why it exists.
3. What data exists
Availability below means "exists on the platform and can in principle be released de-identified through the export route", not "will be handed over on request". Chapter 5 covers how release actually works.
| Data | Granularity | Dimension | Notes |
|---|---|---|---|
| Assistant dialogue | Per message, timestamped, threaded by conversation | Cognomics / Linguomics | The core asset. Includes which mode the conversation was in. |
| Cognitive-level labels | Per student message | Cognomics | Revised Bloom taxonomy, model-assigned. See chapter 9 before you use these as an outcome. |
| Message embeddings | Per message vector | Linguomics | Enables clustering and semantic similarity work without re-embedding. |
| Quizzes, worksheets, assignments | Per item response, with scores | Cognomics | Item content is the teacher's; instruments differ between courses. |
| Forum activity | Per post and reply, with quality and reaction signals | Sociomics | Includes the quality-adjusted scoring components, not only counts. |
| Surveys | Per response | Depends on instrument | In-platform surveys only. New instruments are Case B. |
| Cognitive test battery | Trial level, millisecond timing | Cognomics | Twenty-three research-grade tasks across executive function, processing speed, working memory, episodic memory, fluid reasoning, spatial cognition, verbal ability and metacognition. Trial-level records are never physically deleted. |
| Learner profiling | Per completed instrument | Cognomics | Three open psychometric instruments (Holland RIASEC, IPIP Big Five, OEJTS). Students take them for themselves and may repeat them; all attempts are retained. |
| Wearable daily summaries | Daily values | PhysioNeuromics | Sleep stages and score, heart-rate variability, stress and body battery, activity, resting heart rate. Only within an authorised research project, only for members who connected a device. See chapters 5 and 7. |
| Environmental exposure | Derived daily indicators at institution or county level | Environomics | Computed from public agency monitoring data. Derived indicators can be used as variables; the underlying agency datasets are not redistributed by us — obtain those from the agency directly if you need raw values. |
It is broad rather than uniform. Courses differ in which tools they used, teachers differ in what they required, and participation is voluntary throughout — so coverage is uneven by design and cannot be made even retrospectively. Any design that assumes every student has every modality will not survive contact with the data. Ask for a feasibility count on your specific combination of variables before you commit to a design; that is a cheap question and we would rather answer it early.
4. What you will never get
Not "not yet" — these are settled positions, and asking again does not change them.
- Identified students. Research data is released with random research codes. There is no re-identification key on your side and you must not attempt to construct one.
- Another teacher's course, without them. Course data belongs to the course. Multi-course studies happen through the teachers concerned, not around them.
- Menstrual-cycle and reproductive-health data. Collected only at the service layer for the user's own reference, excluded from every research category, and not authorisable by any project. The reasoning is in chapter 7 of the research governance document: the science is weak and the history of such claims is bad enough that the ethical risk exceeds the academic value.
- Raw continuous physiological signals through the project route. Research projects authorise daily summary values, not minute-level series. Beat-level work exists on the platform but runs as its own instrumented study, not as a data request.
- Student work through the public API. The public API and MCP server expose low-sensitivity data only — public courses, institutions, publications. Dialogue, grades and forum content are never exposed that way, and this boundary is not negotiable per project.
- Data from a student who withdrew. Withdrawal is immediate and excludes the student from subsequent analysis and export. Aggregates already computed and de-identified before withdrawal cannot be retroactively decomposed, and that limit is disclosed to students when they consent.
5. The export route
There is one door. No informal copy, no dataset shared over email, no exception for a collaborator in a hurry — that single-route property is what makes the umbrella protocol workable, so it is enforced rather than encouraged.
What happens on the way out
- Identifiers are replaced with random research codes.
- The export is encrypted and delivered to the requester, and every export is logged.
- The responsible teacher must have signed a data-protection undertaking; a research assistant exporting on their behalf must additionally hold senior teaching-assistant rights and have signed a confidentiality undertaking.
- Research export is bounded in time: records remain exportable for five years from creation, after which the system refuses.
Access tiers for external researchers
Collaborating researchers who are not the course teacher come in through a researcher access agreement covering tiered access, confidentiality, publication review and termination. It is published in the governance centre (Chinese; ask if you need a formal English copy for your institution's file).
Within a research project a principal investigator can view members' authorised daily summaries during the project window. The de-identified research export of that data opens only after the ethics committee approves the corresponding protocol amendment, and it is not open today. If your study depends on exporting wearable data, treat that as an unresolved dependency and talk to us before you write it into a proposal — do not assume the viewing capability implies an export capability.
6. Study designs that fit
| Design | What it needs | Ethics case |
|---|---|---|
| Secondary analysis of existing dialogue — corpus work, trajectory modelling, method comparison | A defined question and a feasibility count. Nothing collected. | A |
| Design-based research on your own course — you change your teaching, the platform records what follows | Your course, a written expectation before you start, and honesty in the write-up about what else changed that semester. | A, if you add no new instrument |
| Pre/post around a designed activity — for example Socratic dialogue with assessment either side | Existing platform instruments. Adding your own test items makes it B. | A or B |
| Cross-course or cross-institution comparison | Each teacher's involvement, and realism about how much differs between courses besides your variable of interest. | A |
| Cognitive testing linked to learning behaviour | Students completing the battery — which is voluntary, so plan for attrition between enrolment and a usable N. | A |
| Physiological work | A research project with per-member authorisation, members who own and connect a device, and the export dependency in chapter 5. | Usually B |
| A new questionnaire or instrument | An amendment, filed with lead time. Then it runs as an in-platform survey. | B |
In-platform surveys are the default for a reason: responses stay inside the existing consent and governance arrangement, are linkable to platform data through the same de-identification route, and do not create a second data controller. An external form breaks all three at once.
7. Consent architecture
Two independent layers, plus a set of participant rights that are implemented in the software rather than promised in a document.
- Platform-level research consent. Versioned, withdrawable, and separate from using the platform at all. Declining it leaves every teaching feature available.
- Per-project authorisation, for research projects involving wearable data. The participant reads the project's stated purpose verbatim, selects which data categories to authorise, and chooses whether history before the join date is included. The full text is hashed at signature for later evidence.
Participant rights, as implemented
- Mirror view. A participant sees exactly the screen the investigator sees. There are no hidden fields.
- Access log. Every view by the investigator is recorded — who, when, what range — and the participant can read that log.
- Withdrawal. One action, immediate effect, no reason required, no consequence. The investigator may not ask why.
- Anonymous alert. A participant who feels pressured but does not want to withdraw can raise a flag without leaving the project and without the investigator knowing who raised it. Once alerts reach the threshold — at least two people and no less than five per cent of active members — the project's viewing access is automatically suspended pending review.
Where an investigator has authority over participants — teacher and student, coach and athlete, commander and subordinate — the project must declare it, and the investigator signs an additional undertaking: data is used for training adjustment and participant care only, never for evaluation, discipline, selection or any personnel decision. Concealing the relationship is a serious violation. Full text and rationale: research governance.
8. The rule that cannot bend
Signing a consent form earns nothing. Declining costs nothing. Withdrawing mid-semester costs nothing. The student attends, uses the assistant and is graded exactly as before; only their data is excluded. No feature, wording, notification or course announcement may suggest otherwise — not "participants get bonus marks", not "the study is part of the course requirements", not even a hint that the teacher would prefer it.
This is not a stylistic preference. Under the power asymmetry between a teacher and the students they grade, consent obtained with a grade attached is not consent, and the consequences in comparable cases have included dismissal, loss of research eligibility, fines and clawback of funding for the institution.
Two consequences worth planning around:
- Recruit in a way that survives scrutiny. An announcement students can ignore is fine. A recruitment pitch delivered by you in your own classroom is a pressure you should design out, not defend later.
- Grading AI use changes your data. Teachers may grade AI use — that is their teaching autonomy. But behaviour produced by a grade incentive is a different phenomenon from spontaneous use, and a paper that does not distinguish them will be asked about it in review.
9. Limitations you should write
Reviewers ask about these, so it is better to raise them yourself. Each of the following is a real property of the data, not a hedge.
Cognitive-level labels are model output
Bloom-level classifications are produced by a language model against a documented prompt, not by trained human coders. They are consistent and useful in aggregate and across time; they are not a ground truth for an individual message, and their agreement with human coding is a property you should either cite from existing work or establish yourself on a sample. State the model and prompt version — both are recorded per event, so you can.
Usage volume is not engagement
Message counts, session counts and word counts are weak proxies and become weaker where a teacher grades participation. Where the platform offers a quality-adjusted signal, prefer it, and say why you chose it.
Participation is voluntary throughout
Students choose whether to use the assistant, whether to take the cognitive battery, whether to connect a device, and whether to consent to research. Every one of those is a selection step. Report the funnel — enrolled, used, completed, consented — rather than only the final N.
Demographic variables are patchy
The platform did not collect biological sex at registration in its early years, and a large share of accounts have no value recorded. Platform-level claims about gender composition are therefore not supportable, and any equity analysis must be scoped to the subset with recorded values, with that subset described honestly. If your study needs a demographic variable, treat obtaining it as Case B and plan accordingly.
Cognitive battery scores are not clinical norms
Composite scores in the summary report are heuristic conversions for feedback to students, not standardised norm-referenced measures. Use the primary trial-level metrics for analysis, and do not describe the battery as equivalent to an intelligence or clinical assessment.
Context and generalisability
The corpus grew up in Taiwanese higher education and is predominantly Chinese-language, in courses taught by teachers who chose to adopt an AI platform. Both facts bound generalisation and both belong in your limitations paragraph, not in a reviewer's report.
10. Methods, citation, acknowledgement
Describing the platform
Note the two things to avoid: describing the platform as any institution's product, and describing it as commercial. Neither is accurate.
Provenance and ethics statement
If your institution or journal needs more than this, ask — a written provenance statement for your specific dataset is something we produce as a matter of course, not a favour.
Conflict of interest
Declare three things, because a reviewer who discovers any of them unaided will assume the worst about all three:
- Whose platform the data came from, and whether that person is an author.
- That platform use is free of charge and no commercial entity funds or profits from it.
- Any grant funding behind the platform capacity your study consumed, and any API costs your institution paid directly.
Citing prior work
If you use an analytic component — cognitive-level classification, retrieval, forum scoring, the cognitive battery — cite the paper describing it rather than citing the website. The list with DOIs is at /practice, and the method notes at /doc/methodology (Chinese) contain the parameter-level detail a methods section needs.
11. Authorship and collaboration
- Co-authorship is the normal arrangement where platform work is part of the contribution — building an instrument, running an extraction pipeline, designing a mechanism the paper depends on. Where the contribution is providing existing data and a provenance statement, acknowledgement is normal and sufficient.
- Decide before the analysis, not before submission. Authorship conversations held late go badly; held early they take five minutes.
- Teachers whose courses supplied the data should be told what is being written about their class, whether or not they are authors. This is a courtesy the platform expects rather than a rule it enforces, and it is what keeps teachers willing to host studies.
- Overlapping submissions. If your question is close to work already under review on the platform, we will tell you — and expect the same in return. Two papers built on the same corpus making the same claim in different venues is a problem for both.
12. Labs and research projects
Physiological and multi-participant studies run through a lab and project structure rather than ad hoc.
- Labs are opened by the platform administrator with a named principal investigator. Write in with your institution and title, the lab name and research direction, and the intended use.
- Research projects are created by a PI, reviewed and activated by the platform, and then recruit members individually. Only a lab PI may create one, deliberately: every project needs one identifiable person answerable for its ethics.
- Participants do not need lab membership. They receive an invitation, read the project purpose, sign, and join — that is the whole path.
- Projects expire. Viewing access closes automatically on the project end date, with no manual step and no way to quietly extend it.
The full lifecycle, the PI's rights and obligations, and a plain-language glossary of the legal and ethical terms involved are in the research governance document (Chinese). It is the authoritative text; this chapter is a summary of it.
13. How to start
Send the question, not a proposal. A paragraph is enough to get a real answer about feasibility, which ethics case you are in, and whether someone is already working on it.
To [email protected]. You will get a direct answer, including "that data does not exist" or "someone is already doing that" where those are true.
Two minutes each: check /practice for whether the question is already answered, and re-read chapter 2 to work out your ethics case. Those two facts change the reply you get more than anything else in the message.
Something in this handbook wrong, out of date, or missing? Write to [email protected] and say which chapter.
All handbooks · AI teaching overview · Publications · Research governance · uedu.tw
Uedu is developed by Chia-Kai Chang, Assistant Professor at the Center for General Education, National Central University, Taiwan, and is adopted by institutions beyond it. This handbook is available in English · 繁體中文 · 简体中文 · 日本語 · 한국어 · Tiếng Việt · Bahasa Indonesia · Bahasa Melayu · ไทย · Türkçe · Deutsch · Français · Español · Português · Italiano · Ελληνικά. The English version is the reference text. The platform interface itself is available in 16 languages.