Start Here: Contributing Study Metadata to CLASSIC
Welcome to CLASSIC
The CLASSIC Portal helps researchers discover longitudinal studies, understand what information each study collected, and determine whether a study may be relevant to their research question.
CLASSIC primarily shares descriptive, study-level metadata rather than individual participant-level records. Contributor metadata may describe:
-
Study identity and purpose
-
Governance and access requirements
-
Cohort and population characteristics
-
Study design and timing
-
Measures and instruments
-
Derived variables and data-availability summaries
This guide explains how to prepare and submit metadata, avoid common curation issues, and review the metadata generated from your study documentation.
New contributor? We strongly recommend scheduling an optional onboarding consultation early in the submission process, before completing the entire template or assembling a large collection of source files.
Schedule a CLASSIC onboarding consultation: [Insert calendar scheduling link]
Submit a question or request support: [Insert service desk or contact link]
1. CLASSIC Metadata Submission Workflow
The general CLASSIC metadata contribution process is:
-
Review this Start Here guide.
-
Schedule an optional onboarding consultation, particularly if this is your first CLASSIC submission.
-
Identify the study materials to be used in preparing metadata.
-
Remove redundant or duplicate source files.
-
Complete Tiers 1–4 of the CLASSIC metadata template.
-
Submit the completed template and supporting documents through the agreed submission location.
-
Sage Bionetworks extracts additional measure- and variable-level metadata where applicable.
-
The contributor reviews the extracted metadata for accuracy and completeness.
-
Sage validates and prepares the metadata for publication.
-
The study is published or updated on the CLASSIC Portal.
Metadata can be improved after the initial publication. A study does not need to have every possible metadata field completed before it can be represented in CLASSIC.
2. What Contributors Are Expected to Complete
The CLASSIC template is organized into metadata tiers.
Tier 1: Study Identity and Discovery
Tier 1 makes the study identifiable and discoverable. It establishes what the study is and how it relates to other studies, cohorts, or projects.
Typical fields include:
-
Study identifier
-
Study name
-
Study acronym or key
-
Study description
-
Short study abstract
-
Grant number
-
Principal investigator information
-
Institution
-
Broad categories of available data
-
Cohort identifiers and names
Tier 2: Governance, Access, and Provenance
Tier 2 explains whether and how researchers may access or reuse the study data.
Typical fields include:
-
Whether the data are externally requestable
-
IRB status and institution
-
Consent type
-
Deidentification approach
-
Access restrictions
-
Data-use conditions
-
Access procedures
-
Access request link
-
Data location
-
Access contact
-
Reviewing body
-
Required access materials
Tier 3: Cohort and Population
Tier 3 describes who participated in the study and how the study population was defined.
Typical fields include:
-
Disease or health focus
-
Ascertainment context
-
Population type
-
Population description
-
Age group
-
Sex and gender eligibility or distributions
-
Race and ethnicity distributions
-
Generational or family relationships
-
Eligibility criteria
-
Inclusion and exclusion characteristics
CLASSIC generally captures these characteristics as aggregate study or cohort summaries rather than individual participant-level records.
Tier 4: Study Design and Time
Tier 4 describes when, where, and how often data were collected.
Typical fields include:
-
Years of data collection
-
Enrollment age or timepoint
-
Administration context
-
Collection modality
-
Recruitment waves
-
Study periods, arms, or components
-
Follow-up timing
-
Rolling enrollment
-
Study start year
-
Burst-design information
-
Observation timestamp availability
Measure- and Variable-Level Metadata
Tier 5 metadata is extracted by Sage Bionetworks from contributor-provided documentation.
This tier may include:
-
Measures and instruments
-
Measured constructs
-
Instrument names and citations
-
Item text
-
Response types
-
Recall periods
-
Derived scores
-
Variable descriptions
-
Data-availability summaries
Contributors are expected to review and validate this extracted information before publication.
3. Complete These Fields First
To reduce rework, complete the template in the following order.
First Priority: Study Identity
Start with:
-
studyID -
studyName -
studyKey -
studyDescription -
studyAbstract -
principalInvestigator -
principalInvestigatorRoles -
principalInvestigatorContact -
institution -
grantNumber
Use the same stable studyID throughout all template tabs.
Second Priority: Study and Cohort Structure
Next, identify the major units represented in the study:
-
Cohorts
-
Recruitment waves
-
Study periods
-
Study arms
-
Follow-up phases
-
Relevant study components
Complete applicable identifiers such as:
-
cohortID -
cohortName -
studyWaveID -
studyPeriodID -
studyArmID -
studyComponentID
Stable identifiers help CLASSIC connect metadata across the different template tiers.
Third Priority: Governance and Availability
Complete the fields that tell researchers whether and how they can obtain the data:
-
externallyRequestable -
accessRestrictions -
dataUseConditions -
accessProcedure -
accessRequestURL -
dataLocation -
accessContact -
accessReviewBody -
accessRequirements
Fourth Priority: Population and Study Design
Complete the study-level or cohort-level summaries that support portal discovery:
-
Population type and description
-
Age group
-
Sex, gender, race, and ethnicity distributions
-
Eligibility criteria
-
Temporal coverage
-
Study waves
-
Follow-up intervals
-
Collection modality
Fifth Priority: Measure and Variable Review
After the measure- and variable-level metadata have been extracted, review:
-
Instrument identification
-
Measured construct
-
Item-to-instrument linkage
-
Response Type
-
Recall Period
-
Derived scores and variables
4. Preparing Source Files
Source documents may include:
-
Codebooks
-
Data dictionaries
-
Instrument documentation
-
Questionnaires
-
Survey forms
-
Variable lists
-
Protocols
-
Data collection manuals
-
Scoring documentation
-
Study websites or publications
Before submitting source files, confirm that each file provides unique or necessary information.
Do Not Submit Redundant Source Files
Do not submit multiple codebooks or documents containing the same metadata unless each file represents a genuinely different version, wave, cohort, instrument, or study component.
The current metadata ingestion process does not automatically deduplicate duplicate content. If identical instrument or variable information appears in multiple source files, it may generate duplicate metadata rows.
For example, avoid submitting:
-
The same codebook in both Excel and PDF format
-
Multiple copies of the same instrument documentation
-
A master codebook and several extracts that repeat the same variables
-
Renamed copies of an unchanged file
-
Multiple study-wave files containing identical instrument definitions without clear wave identifiers
Duplicate metadata can make a study appear to contain more measures or variables than it actually does, increasing the amount of contributor review required.
When Similar Files Should Be Included
Include similar files when they contain meaningful differences, such as:
-
Different study waves
-
Different cohorts
-
Different versions of an instrument
-
Different languages
-
Changes to response options
-
Changes to recall periods
-
Changes to item wording
-
Changes to scoring or derivation methods
Clearly indicate the relevant wave, cohort, period, or version in the filename or accompanying submission notes.
Recommended File-Naming Pattern
Use descriptive filenames whenever possible:
StudyKey_DocumentType_WaveOrCohort_Version_Date
Examples:
-
BCS_Codebook_Wave1_v2_2026-07.xlsx -
BCS_InstrumentGuide_FollowUp_v1.pdf -
BCS_VariableDictionary_FullStudy_2026-07.csv
5. Enter Response Type Information Inline
Response Type information should appear directly in the metadata table on the same row as the relevant item or variable.
Do not place Response Type definitions only in notes below the table, in a separate legend, or elsewhere in the document without connecting them to each applicable item.
For example:
|
instrumentName |
itemText |
responseType |
recallPeriod |
|---|---|---|---|
|
Perceived Stress Scale |
In the last month, how often have you felt nervous or stressed? |
5-point frequency scale: Never to Very Often |
Past month |
This is an interim requirement while improvements to the AI-assisted metadata workflow are being developed as part of the Track 1 AI enhancement.
Entering Response Type inline:
-
Makes the relationship between an item and its response options explicit
-
Reduces incorrect associations during automated extraction
-
Helps reviewers distinguish instruments that use multiple response formats
-
Reduces the likelihood of missing or hallucinated response information
-
Makes contributor validation faster
When multiple items use the same Response Type, repeat the value for each applicable row rather than relying only on a merged cell or a note below the table.
6. Recommended Systematic QC Workflow
The following workflow is recommended for reviewing extracted measure- and variable-level metadata.
Step 1: Filter by Instrument Color
If the review workbook uses colors to distinguish instruments, filter or review one instrument color at a time.
This makes it easier to:
-
Keep related items together
-
Identify rows assigned to the wrong instrument
-
Find duplicate instrument records
-
Compare item wording and response formats within an instrument
-
Review complex codebooks without evaluating the entire workbook at once
Color should be treated as a review aid, not as the only identifier. Instrument names and identifiers should still be populated in the metadata.
Step 2: Review Columns O–AA
In the measure and instrument section of the template, columns O–AA contain core information used to describe and interpret each measure.
Depending on the template version, this area includes fields such as:
-
measureID -
measuredConstruct -
constructDescription -
measureType -
isValidatedMeasure -
measureDescription -
instrumentID -
instrumentName -
instrumentSource -
instrumentCitation -
instrumentVersion -
stimulusResponseMode -
administrationInstructions
Review these columns systematically for each instrument.
Confirm that:
-
The measure and instrument are correctly named
-
The measured construct is appropriate
-
Descriptions reflect the study’s actual implementation
-
Sources and citations refer to the instrument rather than only to the overall study
-
Instrument versions are included when known
-
Administration information is not inferred beyond the available documentation
Step 3: Verify Item-to-Instrument Linkage
Confirm that every item, variable, subscale, or derived score is linked to the correct instrument.
Check for:
-
Items assigned to the wrong instrument
-
Instrument headers interpreted as items
-
Introductory or instructional text interpreted as item text
-
Subscale names interpreted as standalone instruments
-
Repeated items caused by duplicate source files
-
Items that are missing an instrument identifier
-
Items from different instrument versions grouped together
When a source document contains several instruments, carefully verify the beginning and ending rows of each instrument.
Step 4: Review Response Type
For each item, confirm that the Response Type accurately describes how a participant or respondent provided an answer.
Examples include:
-
Yes/No
-
Free text
-
Numeric entry
-
Date
-
Multiple choice
-
Single-select categorical
-
Multi-select categorical
-
Likert agreement scale
-
Frequency scale
-
Severity scale
-
Visual analogue scale
-
Clinician-rated score
-
Sensor-derived observation
Confirm that the response scale is not confused with:
-
The item wording
-
The score derived from the item
-
The recall period
-
Missing-value codes
-
Variable storage type
Where possible, include meaningful response labels, such as:
5-point frequency scale: Never, Almost Never, Sometimes, Fairly Often, Very Often
rather than only:
1–5
Step 5: Review Recall Period
Confirm that the Recall Period is associated with the correct item or instrument.
Examples include:
-
Current
-
Today
-
Past 24 hours
-
Past 7 days
-
Past 2 weeks
-
Past month
-
Past year
-
Since the last visit
-
Lifetime
-
Not applicable
-
Not specified
Do not infer a Recall Period when the source material does not provide one. Use an appropriate null explanation instead.
Step 6: Review Derived Variables
Review variables that are calculated, scored, summarized, or transformed from other values.
Examples include:
-
Total scores
-
Subscale scores
-
Composite scores
-
Standardized scores
-
Change scores
-
Trajectory or slope estimates
-
Risk scores
-
Screening flags
-
Aggregated sensor features
Confirm that:
-
The derived variable is not described as a raw item
-
The score or variable name matches the study documentation
-
The derivation method is summarized accurately
-
The source measure or instrument is identifiable
-
Units and value ranges are included when available
-
The applicable cohort, wave, or timepoint is clear
-
No participant-level values are included in the metadata
Step 7: Conduct a Final Duplicate Review
Before returning the reviewed workbook, filter or sort by:
-
Instrument name
-
Instrument identifier
-
Item text
-
Variable name
-
Study wave
-
Cohort
-
Source file
Look for exact or near-exact duplicate rows. Confirm whether each duplicate represents:
-
A true repeated administration
-
A different wave or cohort
-
A different instrument version
-
A legitimate reuse of an item
-
Redundant source-document content that should be removed
7. Using Null or Missing Values
Do not invent or infer metadata solely to fill an empty field.
When a value cannot be provided, use an approved null reason where supported:
-
Not collected
-
Not available
-
Not yet curated
-
Not shared due to restrictions
-
Unknown
Use these consistently:
Not collected
The study did not collect this information.
Not available
The information may have existed but is not currently available in the submitted materials.
Not yet curated
The information is expected to be added during a later curation phase.
Not shared due to restrictions
The information exists but cannot be shared because of governance, licensing, consent, or other restrictions.
Unknown
The contributor cannot determine whether the information was collected or is available.
If the correct value is missing from a controlled-value list, do not select a misleading alternative. Document the proposed value in the submission ticket or notify the CLASSIC team.
8. Optional 1:1 Onboarding Consultations
CLASSIC offers optional onboarding consultations to help contributors prepare their materials and avoid extensive rework.
Who Should Schedule a Consultation?
A consultation is strongly recommended for:
-
Teams submitting to CLASSIC for the first time
-
Studies with multiple waves or cohorts
-
Studies with large or complex codebooks
-
Studies using multiple versions of the same instrument
-
Studies with substantial longitudinal or derived-variable metadata
-
Teams uncertain about which source documents to submit
-
Teams with specific governance or data-access questions
When Should the Consultation Occur?
Schedule the consultation early—ideally before:
-
Completing the entire metadata template
-
Reformatting all study codebooks
-
Uploading a large set of source documents
-
Combining metadata across waves
-
Making assumptions about controlled vocabulary values
An early discussion can identify the appropriate study structure and source-file strategy before significant work has been completed.
Scheduling Process
-
Open the CLASSIC onboarding calendar: [Insert scheduling link].
-
Select a consultation time.
-
Complete the short intake form.
-
Attach or link representative source materials, when possible.
-
Identify the main questions or decisions needed.
-
Attend the consultation with at least one person familiar with the study documentation.
Intake Questions
The scheduling form will ask for:
-
Contributor name
-
Contributor email
-
Study name
-
Institution
-
Grant or program
-
Expected submission date
-
Number of cohorts or waves
-
Approximate number and type of source files
-
Whether individual-level data are available through another repository
-
Primary questions for the CLASSIC team
-
Links to representative codebooks or study documentation
-
Accessibility accommodations or scheduling notes
Consultation Agenda
A typical consultation may cover:
-
CLASSIC’s purpose and metadata-only scope
-
Which template tiers the contributor should complete
-
The appropriate study, cohort, wave, and timepoint structure
-
Which source files to include or exclude
-
How Sage will extract measure- and variable-level metadata
-
Controlled values and null reasons
-
Governance and access information
-
Expected contributor QC responsibilities
-
Publication and update process
-
Anticipated timeline and next steps
9. Contributor Review Checklist
Before submitting the initial metadata template, review as a contributor that you have:
- I reviewed the Start Here guide.
- I used one stable
studyIDacross all tabs. - I completed the core study identity fields.
- I identified relevant cohorts, waves, periods, or study arms.
- I documented how researchers can request or access the data.
- I reviewed study-level population summaries.
- I reviewed study-design and timing information.
- I removed example rows from the template.
- I removed duplicate or redundant source files.
- I clearly labeled source files by wave, cohort, or version.
- I used null reasons rather than inventing unavailable information.
- I notified the CLASSIC team about any missing controlled values.
Before approving extracted measure and variable metadata:
- I reviewed one instrument or instrument color at a time.
- I reviewed columns O–AA.
- I verified item-to-instrument linkage.
- I reviewed Response Type values.
- I confirmed that Response Type information appears inline.
- I reviewed Recall Period values.
- I reviewed derived variables and scores.
- I looked for duplicate or near-duplicate rows.
- I confirmed that instrument sources and citations are accurate.
- I confirmed that no participant-level values are included.
10. Publication Guide
What Is Published on CLASSIC?
Once metadata have been reviewed and validated, CLASSIC may publish:
-
Study title and description
-
Principal investigator and institution
-
Grant information
-
Study population summaries
-
Study design and temporal coverage
-
Data categories
-
Measures and instruments
-
Variable and derived-score descriptions
-
Aggregate data-availability summaries
-
Governance and access instructions
-
Links to data repositories, study websites, or access-request systems
Publication in CLASSIC does not necessarily mean that the underlying participant-level data are open or hosted by CLASSIC.
Publication Process
-
The contributor submits the required study metadata and source materials.
-
Sage reviews the submission for structure and completeness.
-
Sage extracts additional metadata where applicable.
-
The contributor reviews and validates the extracted information.
-
Sage resolves validation issues and prepares portal records.
-
The study is reviewed in the staging environment.
-
The approved metadata are published to the CLASSIC Portal.
-
Future corrections or enrichments may be submitted as updates.
What Contributors Review Before Publication
Contributors should confirm:
-
Study name and description
-
Principal investigator information
-
Cohort and population summaries
-
Study dates and wave structure
-
Available data categories
-
Instrument and measure identification
-
Governance and access language
-
Data location and request links
-
Any restrictions or limitations that researchers should understand
Updating a Published Study
Metadata may be updated when:
-
A new study wave is completed
-
Additional measures become available
-
Data are deposited in a new repository
-
Access procedures change
-
A codebook is revised
-
An instrument or citation is corrected
-
A new cohort or follow-up period is added
-
Previously unavailable metadata are curated
To request an update, submit a request through [insert service desk link] and identify:
-
The study
-
The affected field, instrument, or file
-
The current value
-
The requested value
-
The reason for the change
-
Any supporting documentation
11. Recruiting Studies for CLASSIC
CLASSIC is actively seeking additional longitudinal studies to improve the discoverability and reuse of research resources.
Studies may be a good fit for CLASSIC when they include:
-
Longitudinal or repeated-measures data
-
Cohort, survey, behavioral, clinical, or digital assessments
-
Measures relevant to health, well-being, development, aging, cognition, or social and environmental factors
-
Data hosted in Synapse or another repository
-
Controlled-access data that can still be described publicly
-
Study documentation that can support measure- and variable-level discovery
Studies do not need to make participant-level data openly available to be represented in CLASSIC. Public metadata can help researchers understand whether a study exists, what it collected, and how to request access.
Interested in Contributing a Study?
To discuss whether your study is a fit:
-
Schedule an introductory consultation: [Insert scheduling link]
-
Contact the CLASSIC team: [Insert contact email]
-
Submit an onboarding request: [Insert service desk link]
12. General Data Availability Statement
The following language may be adapted for publications, grant reports, study websites, or other research outputs that draw on studies identified through CLASSIC.
Standard Data Availability Statement
Metadata describing the study, including available cohorts, study design, measures, instruments, and data-access information, are available through the CLASSIC Portal. The CLASSIC Portal provides study-level and aggregate metadata to support discovery and reuse. Availability of the underlying participant-level data is determined by the contributing study and its designated data repository, governance requirements, consent conditions, and applicable data-use agreements. Researchers should consult the study’s CLASSIC record for current access instructions and contact information.
Statement for Controlled-Access Data
Metadata describing this study are available through the CLASSIC Portal. Participant-level data are not distributed directly through CLASSIC and are subject to the study’s access review process, consent conditions, institutional requirements, and data use restrictions. Researchers should follow the access instructions and links provided in the study’s CLASSIC record.
Statement for Openly Available Data
Metadata describing this study are available through the CLASSIC Portal. The underlying data are available through the repository linked from the study’s CLASSIC record, subject to the repository’s applicable terms of use and citation requirements.
Statement When Only Some Data Are Available
Metadata describing the broader study are available through the CLASSIC Portal. The data currently available for external use represent a subset of the measures, cohorts, waves, or study periods collected by the study. Researchers should review the study’s data-availability summary and access information in CLASSIC to determine which materials are currently requestable.
Recommended Acknowledgment
Researchers who use CLASSIC to identify or interpret a study should acknowledge the CLASSIC Portal and separately cite the original study, dataset, instruments, and repository as required by the contributing study.
Suggested acknowledgment:
Study metadata and data-access information were identified through the CLASSIC Portal, developed and maintained by Sage Bionetworks in collaboration with participating studies.
Contributors may provide additional required acknowledgment or citation language for display with their study record.
13. Frequently Asked Questions
Does CLASSIC require individual participant-level data?
No. CLASSIC is designed to make studies and their available resources discoverable through study-level, cohort-level, measure-level, and aggregate metadata. The underlying participant-level data may remain in a separate open or controlled-access repository.
Must every metadata field be completed?
No. Complete the fields that are applicable and available. Use an appropriate null reason when information is not collected, unavailable, restricted, not yet curated, or unknown.
Can a study be published before all metadata enrichment is complete?
Yes. Core study information can be published and enriched over time. Contributors should prioritize accurate, useful metadata over attempting to complete every possible field before initial publication.
Why should duplicate codebooks be excluded?
The current ingestion process does not automatically deduplicate repeated metadata. Duplicate source content can create duplicate instruments, items, or variables in the extracted metadata.
Why must Response Type be entered inline?
Inline Response Type information makes the relationship between an item and its response options explicit. This reduces extraction errors while improvements to the AI-assisted metadata workflow are under development.
What happens after Sage extracts metadata?
The contributor receives the extracted metadata for systematic review and validation. The metadata are not considered publication-ready until contributor questions and corrections have been addressed.
Can metadata be updated later?
Yes. Published study records can be corrected or enriched when new information, waves, measures, access procedures, or data become available.
14. Getting Help
Schedule an onboarding consultation: [Insert calendar link]
Submit a support request: [Insert service desk link]
Contact the CLASSIC team: [Insert shared email address]
Visit the CLASSIC Portal: https://staging.classicportal.synapse.org/
When requesting support, include the study name, study identifier, current submission stage, and links to the relevant template or source documents.