Internal Working Document · Unlisted · Version 1.0 · June 2026
NeuroFlex Research Design Framework
Four-Cohort Longitudinal Study Architecture, Ground Truth Strategy, Behavioral Measurement Framework, Latent Trait Probe Library, and Daily Checkup Module Design
Executive Summary
This document captures the NeuroFlex research design framework as it has evolved through iterative scientific reasoning. The core shift is from thinking in terms of a consumer wellness tool that might eventually generate research data, to a digital longitudinal research cohort platform — a structured study in which the current Android tool serves as the first data collection instrument for a stratified multi-cohort design.
The four-cohort design solves the "ground truth problem" identified by Professor Žilka: without knowing which participants have confirmed pathological changes, behavioral patterns cannot be scientifically interpreted. By recruiting across four stratified groups — diagnosed, family risk, biomarker-positive, and general population — NeuroFlex generates interpretable longitudinal data without waiting decades for outcomes.
The behavioral measurement layer translates abstract psychological constructs (curiosity, initiative, flexibility, resilience) into concrete, non-invasive digital interaction patterns, distributed across four daily checkup modules (Pulse, Compass, Wander, Haven). These behavioral probes are designed to be genuinely engaging to users while simultaneously encoding latent trait signals for backend scientific analysis.
Contents
The Ground Truth Problem
This is the most important conceptual constraint on NeuroFlex as a research platform, identified in explicit terms by Professor Žilka (Slovak Academy of Sciences) at initial consultation in June 2026.
The problem can be stated simply: if NeuroFlex collects behavioral data from 50,000 users but does not know which users have confirmed pathological changes, confirmed diagnoses, or confirmed biomarker status, then the platform can identify behavioral clusters and trajectory types — but cannot interpret what those clusters mean in relation to Alzheimer's disease.
A pattern without a label is a description, not a biomarker. To establish that a behavioral pattern is associated with pre-clinical Alzheimer's disease, you need a reference label — biomarker status, clinical diagnosis, or eventual outcome — against which the behavioral data can be evaluated. Without this, you cannot distinguish between "behavioral cluster A = higher risk" and "behavioral cluster A = people who work in physically demanding jobs."
The four-cohort design solves this problem by recruiting known-status participants through neurological and psychological clinical channels — allowing behavioral data to be correlated against established biomarker and clinical ground truth from day one, rather than waiting decades for population-level outcome data to accumulate.
Four-Cohort Research Design
NeuroFlex is proposed as a stratified longitudinal research platform recruiting across four cohorts with different levels of established Alzheimer's risk or pathology. Crucially, the same user-facing tool can be used across all cohorts — the stratification exists in the research backend, not in the product experience. This enables blinded or partially-blinded behavioral analysis and prevents users from being stigmatized or behaviorally altered by knowledge of their research group.
Cohort A
Early-Stage Diagnosed
Individuals with confirmed early-stage MCI (Mild Cognitive Impairment) or early Alzheimer's disease diagnosis, referred by neurologists or psychologists. This group serves as the reference anchor: it defines what established cognitive impairment looks like in behavioral trajectory data. The research question for Cohort A is: What behavioral signatures are present in individuals with confirmed disease?
Recruitment via: memory clinics, neurology departments, dementia family support organizations.
Cohort B
Family Members
First-degree relatives (children, siblings) of diagnosed individuals, carrying elevated genetic risk without confirmed pathology. This group is scientifically interesting because: (1) they have intrinsic motivation to participate — concern for their own health; (2) they are likely to show high retention; (3) they represent the population in which early behavioral detection would have the highest preventive value. The question: Are their behavioral profiles measurably closer to Cohort A than to Cohort D?
Recruitment via: neurological practices, family outreach through Cohort A participants.
Cohort C
Biomarker-Positive, Cognitively Normal
Individuals with confirmed amyloid and/or tau pathology on PET or CSF testing, who currently score within normal ranges on standard cognitive assessment. This is arguably the most scientifically valuable group. These individuals have confirmed underlying pathological change but no clinical syndrome — precisely the multi-year pre-symptomatic window during which behavioral signals, if they exist, would first emerge. The question: Do biomarker-positive but cognitively normal individuals show measurable behavioral differences relative to biomarker-negative controls?
Recruitment via: existing research cohorts (ideally DELCODE, Amsterdam UMC); requires advanced biomarker infrastructure available primarily in Scandinavia, Netherlands, Germany.
Cohort D
General Population (Control)
App Store / Google Play downloads by general public. Provides normative behavioral baseline data at scale. Essential for establishing what "normal aging behavioral trajectory" looks like, against which cohort-specific deviations can be characterized. Users can voluntarily opt into anonymized research participation. Future architecture: unique user ID + QR code for data persistence across reinstalls.
Also enables secondary population-level research on healthy behavioral aging, independent of Alzheimer's disease focus.
Research Questions — Ordered by Ambition
NeuroFlex is not designed to prove a single hypothesis. It is designed as a platform for sequential hypothesis testing, where each level of evidence creates the foundation for the next. The following ordering is both scientifically and strategically appropriate — early publication of lower-ambition results builds credibility for later higher-ambition claims.
Feasibility: Can we measure behavioral patterns in a way that is reliable, stable, and consistent over time within individuals? This is the precondition for everything else. Expected sample: N=300–1,000. Timeline: 6–12 months.
Cohort differentiation: Are there measurable behavioral differences between individuals with established cognitive impairment (Cohort A), those at genetic risk (Cohort B), and healthy controls (Cohort D)? Expected sample: N=600–2,400. Timeline: 12–24 months.
Pre-symptomatic signal: Do biomarker-positive but cognitively normal individuals (Cohort C) show behavioral differences relative to biomarker-negative controls, prior to any clinical manifestation? This is the primary novel scientific question. Expected sample: N=200–500 in Cohort C. Timeline: 18–36 months.
Blind classification: Can an AI model — given only behavioral data without any cohort labels — classify participants into their correct group at better-than-chance accuracy? If yes, this constitutes strong convergent evidence for the existence of behavioral biomarkers. Expected sample: all cohorts pooled. Timeline: 24–48 months.
Prediction (longitudinal): Do baseline behavioral patterns predict future diagnostic outcomes — MCI conversion, Alzheimer's diagnosis — in prospective follow-up? This is the most ambitious question and requires multi-year follow-up with a large Cohort D base. Timeline: 5–10 years. This is the "holy grail" — publishable only if sufficient longitudinal depth is achieved.
The reformulation of the primary research mission from "predict Alzheimer's" to "characterise longitudinal behavioural trajectories across populations representing different stages of cognitive health and neurodegenerative risk" is both scientifically more precise and strategically more defensible. It sets achievable near-term milestones while maintaining a clear long-term ambition trajectory.
Sample Size Requirements
Pilot Study — Publication-grade proof of concept
| Cohort | N | Purpose |
|---|---|---|
| A — Diagnosed | 100 | Reference anchor for confirmed disease behavioral signatures |
| B — Family | 150 | Risk-elevated group without confirmed pathology |
| C — Biomarker+ | 100 | Pre-symptomatic gold standard |
| D — General | 300 | Normative baseline |
| Total: 650 | Sufficient for initial feasibility and cohort differentiation publications | |
Serious Research — ML-grade, grant-justifiable
| Cohort | N |
|---|---|
| A — Diagnosed | 300 |
| B — Family | 500 |
| C — Biomarker+ | 300 |
| D — General | 1,000+ |
| Total: ~2,100 | |
Longitudinal Outcome Prediction — "Holy grail" tier
To detect incident MCI or AD cases within a general population cohort at a statistically meaningful rate, total Cohort D needs to be at minimum 10,000+ participants with 2–5 year follow-up. This generates sufficient outcome events (expected MCI conversion rate ~1–2%/year in age-appropriate populations) for predictive modelling. At 50,000+ participants, this approaches the scale of large biobank studies and enables sub-group analyses.
For grant applications, use this language: "Initial research objectives — cohort differentiation and pre-symptomatic signal detection — may be achievable with cohorts of several hundred to several thousand stratified participants. Discovery and validation of predictive behavioral biomarkers for future cognitive decline will likely require longitudinal datasets of tens of thousands of participants observed over multiple years."
Evidence Thresholds — What Counts as Success
There is no single statistical threshold that "proves" a hypothesis. Evidence accumulates progressively. The following framework defines what different levels of evidence would mean for NeuroFlex.
Measurement reliability established. Behavioral metrics show acceptable test-retest consistency within individuals over 4+ weeks. This is a precondition, not a result, but publishable as a methods validation paper.
Cohort differentiation. Statistically significant differences in behavioral trajectories between Cohort A and Cohort D, with a meaningful effect size (Cohen's d > 0.4). Requires replication on a second independent dataset. This is a strong, publishable result that would attract significant scientific attention.
Pre-symptomatic signal (Cohort C vs D). Behavioral differences in biomarker-positive but cognitively normal individuals relative to matched biomarker-negative controls. Classifier performance (AUC > 0.70, i.e., well above 0.5 chance level). This would be a landmark result in digital biomarker research.
Blind AI classification above chance. Model trained only on behavioral data (no cohort labels) outperforms chance classification across the full cohort set. Replicated on held-out data from a separate time period or site.
Prospective prediction. Behavioral patterns at baseline predict MCI or AD onset in prospective follow-up, after controlling for known confounders (age, education, genetic risk, depression). This is the definitive longitudinal biomarker claim. Requires 5–10 year follow-up and a large population sample.
Ground Truth in Practice: Clinical Data and GDPR
What clinical minimum data is needed per participant?
This is the key question to clarify with Professor Žilka at the SAV meeting. Based on published research methodology in comparable studies, the minimum useful clinical dataset per participant likely includes:
Age bracket (5-year bands sufficient: 55–59, 60–64...)
Sex
Education level (years of formal education — proxy for cognitive reserve)
Cohort assignment (A/B/C/D) — does not require name, DOB, or any directly identifiable information
For Cohort A/C: diagnostic category (MCI / early AD / biomarker-positive) — provided by referring clinician, not self-reported
Optional for Cohort B: type of familial relationship to diagnosed person (parent, sibling)
Optional future-state: APOE4 genotype status (requires explicit genetic data consent), biomarker levels at enrollment
Voluntary self-disclosure vs. clinical labeling
There is an important distinction between two approaches to ground truth assignment:
Clinical pathway (Cohorts A, B, C): Participants are referred by neurologists or psychologists who know their status. The clinician assigns the cohort label (as an anonymised code, not name-linked) and the participant consents to data sharing. This is standard in clinical research and does not require the app to collect health data directly — the clinician provides the label through a separate research protocol.
Self-disclosure pathway (Cohort D future): General population users can voluntarily indicate, in a clearly research-framed optional survey: "Have you received a diagnosis related to cognitive health since joining NeuroFlex?" This generates prospective outcome labels over time. GDPR-sensitive. Requires: informed consent, explicit opt-in, ethics approval, and clear limitation to research use.
In Phase 1 (MVP), collect no health data. Behavioural data only. Clinical labeling happens through a parallel research protocol administered by the clinical partner, not through the app. This keeps the app legally simple and focuses development on what matters: the behavioral data collection engine.
Persistent User Identity (future architecture)
For longitudinal research validity, user data must be persistent across reinstalls, device changes, and app updates. Proposed architecture: unique anonymized research ID assigned at registration + QR code for account recovery without requiring personal email or phone number. This maintains longitudinal data integrity while preserving anonymity. Implementation target: Phase 2.
Behavioral Measurement Framework
7.1 Why Proxy Signals — Not Direct Measurement
Abstract traits like curiosity, resilience, or cognitive flexibility cannot be asked about directly — "Rate your curiosity from 0 to 10 today" produces noise, not signal. Instead, NeuroFlex captures behavioral proxies: observable interaction patterns from which latent trait states can be inferred by machine learning models.
This approach is methodologically analogous to how personality psychology operationalises abstract constructs (the Big Five personality traits are not measured by asking "how neurotic are you?" — they are inferred from patterns of responses to seemingly unrelated situational questions). The key difference is that NeuroFlex captures these patterns continuously and in context, rather than in a single cross-sectional questionnaire administration.
7.2 Twelve Latent Traits and Their Proxy Signals
1. Curiosity
Proxy signals: novel content engagement rate; feature exploration diversity; optional learning content access; "surprise me" choices; deviation from established activity patterns.
2. Initiative
Proxy signals: ratio of spontaneous (notification-independent) session opens to notification-triggered opens; voluntary task completion beyond required minimum; self-initiated social contact planning; temporal regularity of self-directed engagement.
3. Cognitive Flexibility
Proxy signals: activity type switching frequency; performance consistency across diverse cognitive task formats; adaptation speed after routine disruption; willingness to try new task variants.
4. Persistence
Proxy signals: task completion rate; retry rate following failure; session continuation after first difficulty; sustained engagement on longer-format activities; goal-directed behavior maintenance.
5. Resilience
Proxy signals: engagement recovery speed following low-activity periods; performance restoration rate after detected poor performance; re-engagement pattern after life disruptions; stability of engagement despite setbacks.
6. Courage / Uncertainty Tolerance
Proxy signals: preference for uncertain rewards over guaranteed smaller rewards; novel activity acceptance rate; engagement with unfamiliar task formats; willingness to engage with socially uncertain scenarios.
7. Confidence / Self-Efficacy
Proxy signals: self-reported certainty ratings after task completion; calibration between confidence and accuracy (overconfidence vs. underconfidence); decision speed; second-guessing patterns.
8. Social Activation
Proxy signals: completion rate of social/Connect module tasks; spontaneous vs. prompted social contact reporting; preference for social vs. solitary activities in binary choices; social interaction quality self-assessment trends.
9. Future Orientation
Proxy signals: planning horizon (today / this week / next month); temporal preference in activity choices; optimism about upcoming period; engagement with goal-setting features.
10. Self-Awareness
Proxy signals: self-prediction accuracy (morning prediction vs. evening actual report); ability to detect own state changes; accuracy of own performance estimation; self-monitoring engagement quality.
11. Decisiveness
Proxy signals: decision latency (time to choice in binary selection tasks); decision regret indicators; avoidance of choice; delegation to "default" or "random" options.
12. Routine Preference
Proxy signals: consistency of daily engagement timing; preference for familiar vs. novel content when offered equivalent choice; response to unexpected schedule disruption; entropy of daily interaction pattern.
7.3 What Matters: Trajectory, Not Absolute Score
The absolute score on any trait measure is less scientifically interesting than its trajectory over time. Two participants with identical average Curiosity scores can have dramatically different trajectories — one stable, one declining. Furthermore, increased variability (higher day-to-day inconsistency in trait expression) may itself be a meaningful signal, independent of direction. This is why years of daily data — not a single assessment — are the core scientific asset of NeuroFlex.
Daily Checkup Architecture — Four Modules
The four daily checkup modules are the primary behavioral data collection interface. Each module runs once per day, contains 1–3 questions or micro-tasks, and is presented in a naturalistic, engaging format that does not feel clinical. Questions rotate to prevent habituation and learning effects. Backend tagging maps each question to one or more latent trait domains.
Pulse — Inner State
Self-assessment of current mood, energy, focus, and body awareness. The most direct window into subjective wellbeing and self-monitoring capacity.
Timing: morning (captures starting state). 2–3 short questions. Includes a visual mood input and one contextual self-prediction question.
Compass — Decision & Risk
Hypothetical decision scenarios and preference questions probing risk tolerance, future orientation, decisiveness, initiative, and uncertainty tolerance.
Timing: mid-day. 1–2 scenario questions. No right/wrong answers. Presented as "what would you do?" or "which sounds better?"
Wander — Curiosity & Flexibility
Novelty-seeking, adaptability, and openness to change. Questions probe how users relate to new experiences, learning, disruption, and routine versus variety.
Timing: afternoon. 1–2 questions. May include a micro-exploration task (e.g., "discover one new fact today").
Haven — Social & Resilience
Social activation, relationship quality, support-seeking, and resilience. Probes how users navigate interpersonal connection, response to setbacks, and recovery patterns.
Timing: evening. 1–2 questions or a brief social micro-task report. May include "Did you connect with someone today?" type prompts.
Questions rotate across modules and are never presented as a battery. The user experiences 4–6 micro-interactions per day across the four modules — each taking under 30 seconds. No day feels identical. Backend tagging links each response to latent traits regardless of which module it appears in. Over 365 days, this generates approximately 1,500–2,000 individual behavioral data points per user, across 12 trait dimensions.
Behavioral Probe Library — Sample Questions
The following examples illustrate the probe design philosophy: questions appear as natural preference or scenario questions; backend tagging encodes the scientific intent. Users are never told which trait is being assessed.
Curiosity — Novel Content Choice
Which headline would you open?
latent_traits: [curiosity, novelty_seeking]Uncertainty Tolerance — Reward Preference
Which would you choose?
latent_traits: [courage, uncertainty_tolerance, decisiveness]Persistence — Task Difficulty Response
When something becomes harder than expected, you usually:
latent_traits: [persistence, resilience]Initiative — Spontaneous Action
When you have unexpected free time, you usually:
latent_traits: [initiative, routine_preference]Confidence Calibration — Post-task
How confident are you in your answer?
latent_traits: [confidence, self_awareness] — compared against actual performance to compute calibration_indexFlexibility — Disruption Response
Your usual plan changes unexpectedly today. You:
latent_traits: [flexibility, resilience, routine_preference]Self-Prediction (Morning)
Do you think you'll exercise today?
[Evening follow-up] Did you exercise today?
derived_metric: self_prediction_accuracy — tracked longitudinally per userSocial Activation — Connection Initiative
If you think of someone you haven't spoken to in a while, you usually:
latent_traits: [social_activation, initiative]Slovak Context and International Research Strategy
Professor Žilka's Key Observation
The core constraint identified at the SAV consultation is that Slovak clinical diagnostic infrastructure — particularly for biomarker-positive pre-symptomatic individuals (Cohort C) — is substantially less developed than in Scandinavian countries, the Netherlands, or Germany. PET amyloid scanning, CSF biomarker programs, and systematic pre-symptomatic screening exist at significantly higher rates in countries like Sweden, Finland, Denmark, and the Netherlands. This makes Slovak-only recruitment of Cohort C difficult.
This does not prevent NeuroFlex from being a Slovak startup. It means that the research collaboration strategy must be international from day one. The platform and its first Android tool are developed in Slovakia; the research cohorts are international. This model is standard in European digital health research — it is exactly the structure used by RADAR-AD, LETHE, and DigiAD, all of which had primary development in one country and clinical recruitment sites across multiple EU countries.
Recommended Partnership Structure
NeuroFlex s.r.o. — platform development, AI/ML, data infrastructure, project coordination. Slovak Academy of Sciences (SAV) as scientific advisory and local ethics approval partner. KInIT Bratislava as computational research partner. Comenius University neurology/psychology for local Cohort A/B pilot.
DZNE / DELCODE — Cohort C clinical ground truth. German biomarker infrastructure, existing SCD cohort, DZNE ethics framework. Amsterdam UMC (Amsterdam Dementia Cohort) as second Cohort C/A site. FINGER/World Wide FINGERS for multidomain intervention context (Cohort D at scale).
Questions for the SAV Meeting
What is the minimum clinical dataset required per participant for the behavioral data to be scientifically interpretable?
Are there existing Slovak or Czech cohorts (Memory Clinic Bratislava, Masaryk University Brno, Czech National Cohort) to which NeuroFlex could be attached as a digital behavioral monitoring module?
Is SAV's Institute of Neuroimmunology the appropriate scientific partner for a formal collaboration agreement, and what would that process look like?
Which specific grant mechanisms does Professor Žilka recommend as the most realistic near-term targets for NeuroFlex? (See the Grant Landscape document for the currently identified options.)