Internal Working Document · Unlisted · Version 1.0 · June 2026

NeuroFlex Research Design Framework

Four-Cohort Longitudinal Study Architecture, Ground Truth Strategy, Behavioral Measurement Framework, Latent Trait Probe Library, and Daily Checkup Module Design

Executive Summary

This document captures the NeuroFlex research design framework as it has evolved through iterative scientific reasoning. The core shift is from thinking in terms of a consumer wellness tool that might eventually generate research data, to a digital longitudinal research cohort platform — a structured study in which the current Android tool serves as the first data collection instrument for a stratified multi-cohort design.

The four-cohort design solves the "ground truth problem" identified by Professor Žilka: without knowing which participants have confirmed pathological changes, behavioral patterns cannot be scientifically interpreted. By recruiting across four stratified groups — diagnosed, family risk, biomarker-positive, and general population — NeuroFlex generates interpretable longitudinal data without waiting decades for outcomes.

The behavioral measurement layer translates abstract psychological constructs (curiosity, initiative, flexibility, resilience) into concrete, non-invasive digital interaction patterns, distributed across four daily checkup modules (Pulse, Compass, Wander, Haven). These behavioral probes are designed to be genuinely engaging to users while simultaneously encoding latent trait signals for backend scientific analysis.


Section 1

The Ground Truth Problem

This is the most important conceptual constraint on NeuroFlex as a research platform, identified in explicit terms by Professor Žilka (Slovak Academy of Sciences) at initial consultation in June 2026.

The problem can be stated simply: if NeuroFlex collects behavioral data from 50,000 users but does not know which users have confirmed pathological changes, confirmed diagnoses, or confirmed biomarker status, then the platform can identify behavioral clusters and trajectory types — but cannot interpret what those clusters mean in relation to Alzheimer's disease.

A pattern without a label is a description, not a biomarker. To establish that a behavioral pattern is associated with pre-clinical Alzheimer's disease, you need a reference label — biomarker status, clinical diagnosis, or eventual outcome — against which the behavioral data can be evaluated. Without this, you cannot distinguish between "behavioral cluster A = higher risk" and "behavioral cluster A = people who work in physically demanding jobs."

The four-cohort design solves this problem by recruiting known-status participants through neurological and psychological clinical channels — allowing behavioral data to be correlated against established biomarker and clinical ground truth from day one, rather than waiting decades for population-level outcome data to accumulate.


Section 2

Four-Cohort Research Design

NeuroFlex is proposed as a stratified longitudinal research platform recruiting across four cohorts with different levels of established Alzheimer's risk or pathology. Crucially, the same user-facing tool can be used across all cohorts — the stratification exists in the research backend, not in the product experience. This enables blinded or partially-blinded behavioral analysis and prevents users from being stigmatized or behaviorally altered by knowledge of their research group.

Cohort A

Early-Stage Diagnosed

Individuals with confirmed early-stage MCI (Mild Cognitive Impairment) or early Alzheimer's disease diagnosis, referred by neurologists or psychologists. This group serves as the reference anchor: it defines what established cognitive impairment looks like in behavioral trajectory data. The research question for Cohort A is: What behavioral signatures are present in individuals with confirmed disease?

Recruitment via: memory clinics, neurology departments, dementia family support organizations.

Cohort B

Family Members

First-degree relatives (children, siblings) of diagnosed individuals, carrying elevated genetic risk without confirmed pathology. This group is scientifically interesting because: (1) they have intrinsic motivation to participate — concern for their own health; (2) they are likely to show high retention; (3) they represent the population in which early behavioral detection would have the highest preventive value. The question: Are their behavioral profiles measurably closer to Cohort A than to Cohort D?

Recruitment via: neurological practices, family outreach through Cohort A participants.

Cohort C

Biomarker-Positive, Cognitively Normal

Individuals with confirmed amyloid and/or tau pathology on PET or CSF testing, who currently score within normal ranges on standard cognitive assessment. This is arguably the most scientifically valuable group. These individuals have confirmed underlying pathological change but no clinical syndrome — precisely the multi-year pre-symptomatic window during which behavioral signals, if they exist, would first emerge. The question: Do biomarker-positive but cognitively normal individuals show measurable behavioral differences relative to biomarker-negative controls?

Recruitment via: existing research cohorts (ideally DELCODE, Amsterdam UMC); requires advanced biomarker infrastructure available primarily in Scandinavia, Netherlands, Germany.

Cohort D

General Population (Control)

App Store / Google Play downloads by general public. Provides normative behavioral baseline data at scale. Essential for establishing what "normal aging behavioral trajectory" looks like, against which cohort-specific deviations can be characterized. Users can voluntarily opt into anonymized research participation. Future architecture: unique user ID + QR code for data persistence across reinstalls.

Also enables secondary population-level research on healthy behavioral aging, independent of Alzheimer's disease focus.


Section 3

Research Questions — Ordered by Ambition

NeuroFlex is not designed to prove a single hypothesis. It is designed as a platform for sequential hypothesis testing, where each level of evidence creates the foundation for the next. The following ordering is both scientifically and strategically appropriate — early publication of lower-ambition results builds credibility for later higher-ambition claims.

1

Feasibility: Can we measure behavioral patterns in a way that is reliable, stable, and consistent over time within individuals? This is the precondition for everything else. Expected sample: N=300–1,000. Timeline: 6–12 months.

2

Cohort differentiation: Are there measurable behavioral differences between individuals with established cognitive impairment (Cohort A), those at genetic risk (Cohort B), and healthy controls (Cohort D)? Expected sample: N=600–2,400. Timeline: 12–24 months.

3

Pre-symptomatic signal: Do biomarker-positive but cognitively normal individuals (Cohort C) show behavioral differences relative to biomarker-negative controls, prior to any clinical manifestation? This is the primary novel scientific question. Expected sample: N=200–500 in Cohort C. Timeline: 18–36 months.

4

Blind classification: Can an AI model — given only behavioral data without any cohort labels — classify participants into their correct group at better-than-chance accuracy? If yes, this constitutes strong convergent evidence for the existence of behavioral biomarkers. Expected sample: all cohorts pooled. Timeline: 24–48 months.

5

Prediction (longitudinal): Do baseline behavioral patterns predict future diagnostic outcomes — MCI conversion, Alzheimer's diagnosis — in prospective follow-up? This is the most ambitious question and requires multi-year follow-up with a large Cohort D base. Timeline: 5–10 years. This is the "holy grail" — publishable only if sufficient longitudinal depth is achieved.

The reformulation of the primary research mission from "predict Alzheimer's" to "characterise longitudinal behavioural trajectories across populations representing different stages of cognitive health and neurodegenerative risk" is both scientifically more precise and strategically more defensible. It sets achievable near-term milestones while maintaining a clear long-term ambition trajectory.


Section 4

Sample Size Requirements

Pilot Study — Publication-grade proof of concept

CohortNPurpose
A — Diagnosed100Reference anchor for confirmed disease behavioral signatures
B — Family150Risk-elevated group without confirmed pathology
C — Biomarker+100Pre-symptomatic gold standard
D — General300Normative baseline
Total: 650Sufficient for initial feasibility and cohort differentiation publications

Serious Research — ML-grade, grant-justifiable

CohortN
A — Diagnosed300
B — Family500
C — Biomarker+300
D — General1,000+
Total: ~2,100

Longitudinal Outcome Prediction — "Holy grail" tier

To detect incident MCI or AD cases within a general population cohort at a statistically meaningful rate, total Cohort D needs to be at minimum 10,000+ participants with 2–5 year follow-up. This generates sufficient outcome events (expected MCI conversion rate ~1–2%/year in age-appropriate populations) for predictive modelling. At 50,000+ participants, this approaches the scale of large biobank studies and enables sub-group analyses.

For grant applications, use this language: "Initial research objectives — cohort differentiation and pre-symptomatic signal detection — may be achievable with cohorts of several hundred to several thousand stratified participants. Discovery and validation of predictive behavioral biomarkers for future cognitive decline will likely require longitudinal datasets of tens of thousands of participants observed over multiple years."


Section 5

Evidence Thresholds — What Counts as Success

There is no single statistical threshold that "proves" a hypothesis. Evidence accumulates progressively. The following framework defines what different levels of evidence would mean for NeuroFlex.

L1

Measurement reliability established. Behavioral metrics show acceptable test-retest consistency within individuals over 4+ weeks. This is a precondition, not a result, but publishable as a methods validation paper.

L2

Cohort differentiation. Statistically significant differences in behavioral trajectories between Cohort A and Cohort D, with a meaningful effect size (Cohen's d > 0.4). Requires replication on a second independent dataset. This is a strong, publishable result that would attract significant scientific attention.

L3

Pre-symptomatic signal (Cohort C vs D). Behavioral differences in biomarker-positive but cognitively normal individuals relative to matched biomarker-negative controls. Classifier performance (AUC > 0.70, i.e., well above 0.5 chance level). This would be a landmark result in digital biomarker research.

L4

Blind AI classification above chance. Model trained only on behavioral data (no cohort labels) outperforms chance classification across the full cohort set. Replicated on held-out data from a separate time period or site.

L5

Prospective prediction. Behavioral patterns at baseline predict MCI or AD onset in prospective follow-up, after controlling for known confounders (age, education, genetic risk, depression). This is the definitive longitudinal biomarker claim. Requires 5–10 year follow-up and a large population sample.


Section 6

Ground Truth in Practice: Clinical Data and GDPR

What clinical minimum data is needed per participant?

This is the key question to clarify with Professor Žilka at the SAV meeting. Based on published research methodology in comparable studies, the minimum useful clinical dataset per participant likely includes:

1

Age bracket (5-year bands sufficient: 55–59, 60–64...)

2

Sex

3

Education level (years of formal education — proxy for cognitive reserve)

4

Cohort assignment (A/B/C/D) — does not require name, DOB, or any directly identifiable information

5

For Cohort A/C: diagnostic category (MCI / early AD / biomarker-positive) — provided by referring clinician, not self-reported

6

Optional for Cohort B: type of familial relationship to diagnosed person (parent, sibling)

7

Optional future-state: APOE4 genotype status (requires explicit genetic data consent), biomarker levels at enrollment

Voluntary self-disclosure vs. clinical labeling

There is an important distinction between two approaches to ground truth assignment:

Clinical pathway (Cohorts A, B, C): Participants are referred by neurologists or psychologists who know their status. The clinician assigns the cohort label (as an anonymised code, not name-linked) and the participant consents to data sharing. This is standard in clinical research and does not require the app to collect health data directly — the clinician provides the label through a separate research protocol.

Self-disclosure pathway (Cohort D future): General population users can voluntarily indicate, in a clearly research-framed optional survey: "Have you received a diagnosis related to cognitive health since joining NeuroFlex?" This generates prospective outcome labels over time. GDPR-sensitive. Requires: informed consent, explicit opt-in, ethics approval, and clear limitation to research use.

In Phase 1 (MVP), collect no health data. Behavioural data only. Clinical labeling happens through a parallel research protocol administered by the clinical partner, not through the app. This keeps the app legally simple and focuses development on what matters: the behavioral data collection engine.

Persistent User Identity (future architecture)

For longitudinal research validity, user data must be persistent across reinstalls, device changes, and app updates. Proposed architecture: unique anonymized research ID assigned at registration + QR code for account recovery without requiring personal email or phone number. This maintains longitudinal data integrity while preserving anonymity. Implementation target: Phase 2.


Section 7

Behavioral Measurement Framework

7.1 Why Proxy Signals — Not Direct Measurement

Abstract traits like curiosity, resilience, or cognitive flexibility cannot be asked about directly — "Rate your curiosity from 0 to 10 today" produces noise, not signal. Instead, NeuroFlex captures behavioral proxies: observable interaction patterns from which latent trait states can be inferred by machine learning models.

This approach is methodologically analogous to how personality psychology operationalises abstract constructs (the Big Five personality traits are not measured by asking "how neurotic are you?" — they are inferred from patterns of responses to seemingly unrelated situational questions). The key difference is that NeuroFlex captures these patterns continuously and in context, rather than in a single cross-sectional questionnaire administration.

7.2 Twelve Latent Traits and Their Proxy Signals

1. Curiosity

wander

Proxy signals: novel content engagement rate; feature exploration diversity; optional learning content access; "surprise me" choices; deviation from established activity patterns.

2. Initiative

compass

Proxy signals: ratio of spontaneous (notification-independent) session opens to notification-triggered opens; voluntary task completion beyond required minimum; self-initiated social contact planning; temporal regularity of self-directed engagement.

3. Cognitive Flexibility

wander

Proxy signals: activity type switching frequency; performance consistency across diverse cognitive task formats; adaptation speed after routine disruption; willingness to try new task variants.

4. Persistence

pulsecompass

Proxy signals: task completion rate; retry rate following failure; session continuation after first difficulty; sustained engagement on longer-format activities; goal-directed behavior maintenance.

5. Resilience

haven

Proxy signals: engagement recovery speed following low-activity periods; performance restoration rate after detected poor performance; re-engagement pattern after life disruptions; stability of engagement despite setbacks.

6. Courage / Uncertainty Tolerance

compass

Proxy signals: preference for uncertain rewards over guaranteed smaller rewards; novel activity acceptance rate; engagement with unfamiliar task formats; willingness to engage with socially uncertain scenarios.

7. Confidence / Self-Efficacy

pulse

Proxy signals: self-reported certainty ratings after task completion; calibration between confidence and accuracy (overconfidence vs. underconfidence); decision speed; second-guessing patterns.

8. Social Activation

haven

Proxy signals: completion rate of social/Connect module tasks; spontaneous vs. prompted social contact reporting; preference for social vs. solitary activities in binary choices; social interaction quality self-assessment trends.

9. Future Orientation

compass

Proxy signals: planning horizon (today / this week / next month); temporal preference in activity choices; optimism about upcoming period; engagement with goal-setting features.

10. Self-Awareness

pulse

Proxy signals: self-prediction accuracy (morning prediction vs. evening actual report); ability to detect own state changes; accuracy of own performance estimation; self-monitoring engagement quality.

11. Decisiveness

compass

Proxy signals: decision latency (time to choice in binary selection tasks); decision regret indicators; avoidance of choice; delegation to "default" or "random" options.

12. Routine Preference

pulsewander

Proxy signals: consistency of daily engagement timing; preference for familiar vs. novel content when offered equivalent choice; response to unexpected schedule disruption; entropy of daily interaction pattern.

7.3 What Matters: Trajectory, Not Absolute Score

The absolute score on any trait measure is less scientifically interesting than its trajectory over time. Two participants with identical average Curiosity scores can have dramatically different trajectories — one stable, one declining. Furthermore, increased variability (higher day-to-day inconsistency in trait expression) may itself be a meaningful signal, independent of direction. This is why years of daily data — not a single assessment — are the core scientific asset of NeuroFlex.


Section 8

Daily Checkup Architecture — Four Modules

The four daily checkup modules are the primary behavioral data collection interface. Each module runs once per day, contains 1–3 questions or micro-tasks, and is presented in a naturalistic, engaging format that does not feel clinical. Questions rotate to prevent habituation and learning effects. Backend tagging maps each question to one or more latent trait domains.

❤️

Pulse — Inner State

Self-assessment of current mood, energy, focus, and body awareness. The most direct window into subjective wellbeing and self-monitoring capacity.

Timing: morning (captures starting state). 2–3 short questions. Includes a visual mood input and one contextual self-prediction question.

self_awareness confidence persistence routine_preference
🧭

Compass — Decision & Risk

Hypothetical decision scenarios and preference questions probing risk tolerance, future orientation, decisiveness, initiative, and uncertainty tolerance.

Timing: mid-day. 1–2 scenario questions. No right/wrong answers. Presented as "what would you do?" or "which sounds better?"

initiative courage decisiveness future_orientation persistence
🌸

Wander — Curiosity & Flexibility

Novelty-seeking, adaptability, and openness to change. Questions probe how users relate to new experiences, learning, disruption, and routine versus variety.

Timing: afternoon. 1–2 questions. May include a micro-exploration task (e.g., "discover one new fact today").

curiosity flexibility persistence routine_preference
🌳

Haven — Social & Resilience

Social activation, relationship quality, support-seeking, and resilience. Probes how users navigate interpersonal connection, response to setbacks, and recovery patterns.

Timing: evening. 1–2 questions or a brief social micro-task report. May include "Did you connect with someone today?" type prompts.

social_activation resilience persistence self_awareness

Questions rotate across modules and are never presented as a battery. The user experiences 4–6 micro-interactions per day across the four modules — each taking under 30 seconds. No day feels identical. Backend tagging links each response to latent traits regardless of which module it appears in. Over 365 days, this generates approximately 1,500–2,000 individual behavioral data points per user, across 12 trait dimensions.


Section 9

Behavioral Probe Library — Sample Questions

The following examples illustrate the probe design philosophy: questions appear as natural preference or scenario questions; backend tagging encodes the scientific intent. Users are never told which trait is being assessed.

Curiosity — Novel Content Choice

wander

Which headline would you open?

→ "5 things you never knew about the ocean" → "Today's news summary" → "Tips for better sleep" → "I wouldn't open any of these"
latent_traits: [curiosity, novelty_seeking]

Uncertainty Tolerance — Reward Preference

compass

Which would you choose?

→ A small reward, guaranteed → A larger reward, uncertain
latent_traits: [courage, uncertainty_tolerance, decisiveness]

Persistence — Task Difficulty Response

pulsecompass

When something becomes harder than expected, you usually:

→ Keep going until finished → Take a short break and continue → Return to it later → Leave it unfinished
latent_traits: [persistence, resilience]

Initiative — Spontaneous Action

compass

When you have unexpected free time, you usually:

→ Start doing something straightaway → Wait for something to come up → Continue a routine activity → Just rest
latent_traits: [initiative, routine_preference]

Confidence Calibration — Post-task

pulse

How confident are you in your answer?

→ Very confident → Somewhat confident → Unsure → Guessing
latent_traits: [confidence, self_awareness] — compared against actual performance to compute calibration_index

Flexibility — Disruption Response

wander

Your usual plan changes unexpectedly today. You:

→ Adapt easily → Need some time to adjust → Feel frustrated → Prefer to cancel the whole thing
latent_traits: [flexibility, resilience, routine_preference]

Self-Prediction (Morning)

pulse

Do you think you'll exercise today?

→ Yes / Maybe / No

[Evening follow-up] Did you exercise today?

→ Yes / No / Partially
derived_metric: self_prediction_accuracy — tracked longitudinally per user

Social Activation — Connection Initiative

haven

If you think of someone you haven't spoken to in a while, you usually:

→ Contact them → Think about it but don't reach out → Wait for them to contact you → Don't think about it much
latent_traits: [social_activation, initiative]

Section 10

Slovak Context and International Research Strategy

Professor Žilka's Key Observation

The core constraint identified at the SAV consultation is that Slovak clinical diagnostic infrastructure — particularly for biomarker-positive pre-symptomatic individuals (Cohort C) — is substantially less developed than in Scandinavian countries, the Netherlands, or Germany. PET amyloid scanning, CSF biomarker programs, and systematic pre-symptomatic screening exist at significantly higher rates in countries like Sweden, Finland, Denmark, and the Netherlands. This makes Slovak-only recruitment of Cohort C difficult.

This does not prevent NeuroFlex from being a Slovak startup. It means that the research collaboration strategy must be international from day one. The platform and its first Android tool are developed in Slovakia; the research cohorts are international. This model is standard in European digital health research — it is exactly the structure used by RADAR-AD, LETHE, and DigiAD, all of which had primary development in one country and clinical recruitment sites across multiple EU countries.

Recommended Partnership Structure

SK

NeuroFlex s.r.o. — platform development, AI/ML, data infrastructure, project coordination. Slovak Academy of Sciences (SAV) as scientific advisory and local ethics approval partner. KInIT Bratislava as computational research partner. Comenius University neurology/psychology for local Cohort A/B pilot.

EU

DZNE / DELCODE — Cohort C clinical ground truth. German biomarker infrastructure, existing SCD cohort, DZNE ethics framework. Amsterdam UMC (Amsterdam Dementia Cohort) as second Cohort C/A site. FINGER/World Wide FINGERS for multidomain intervention context (Cohort D at scale).

Questions for the SAV Meeting

What is the minimum clinical dataset required per participant for the behavioral data to be scientifically interpretable?
Are there existing Slovak or Czech cohorts (Memory Clinic Bratislava, Masaryk University Brno, Czech National Cohort) to which NeuroFlex could be attached as a digital behavioral monitoring module?
Is SAV's Institute of Neuroimmunology the appropriate scientific partner for a formal collaboration agreement, and what would that process look like?
Which specific grant mechanisms does Professor Žilka recommend as the most realistic near-term targets for NeuroFlex? (See the Grant Landscape document for the currently identified options.)