1. Data Sources
Every metric in ThinkKits is derived from verified, publicly available federal and state datasets. We never estimate or infer data points — if it's not in a federal source, we don't show it.
| Source | Agency | Primary Use | Frequency |
|---|---|---|---|
| NCES Common Core of Data | Dept. of Education | School identification, enrollment, demographics | Annual |
| Census SAIPE | Census Bureau | Poverty estimates for Title I formulas | Annual |
| USAC E-Rate | FCC / USAC | Technology spending, vendor commitments | Quarterly |
| CRDC | Dept. of Ed / OCR | Discipline, AP enrollment, staffing, expenditures | Biennial |
| EDFacts | Dept. of Education | Assessment proficiency, graduation rates | Annual |
| F-33 Finance Survey | NCES / Census | Revenue, per-pupil expenditure | Annual |
| Title I/II/III/IV-A | Dept. of Education | Federal program allocations | Annual |
| IDEA Part B | Dept. of Ed / OSEP | Special education funding | Annual |
| DOE Teacher Shortage Areas | Dept. of Education | Shortage Map — teacher staffing gaps | Annual |
See our full Data Source Catalog for complete source documentation, or review our Data Quality & Validation report for coverage metrics and accuracy cross-checks.
2. School Health Score
The School Health Score is a composite 0-100 metric designed to give a quick, at-a-glance assessment of a school's resource adequacy and student outcomes. It combines three weighted dimensions:
2.1 Dimensions
| Dimension | Weight | Inputs |
|---|---|---|
| Funding Adequacy | 40% | Per-pupil expenditure vs. state median, Title I allocation ratio, E-Rate discount rate |
| Demographic Equity | 30% | FRL rate, poverty concentration (SAIPE), student-teacher ratio vs. state median |
| Performance Indicators | 30% | EDFacts proficiency rates (math + reading), chronic absenteeism rate (CRDC), graduation rate (where applicable) |
2.2 Calculation
Each dimension is normalized to a 0-100 scale using min-max normalization within the school's state. The composite score is a weighted average:
Schools with missing data in one dimension receive a weighted score from the remaining dimensions. If more than one dimension is missing, no Health Score is calculated.
The Health Score is a screening tool, not a definitive judgment of school quality. It is designed to surface schools that may benefit from additional resources — not to rank schools against each other. Always combine with qualitative context.
3. Peer Matching Algorithm
The Peer Comparison tool identifies schools that are statistically similar to a selected school. Peer matching uses a nearest-neighbor approach across the following features:
3.1 Matching Features
- Enrollment: Total student count (log-transformed to handle scale)
- FRL Rate: Free/reduced lunch percentage (poverty proxy)
- Locale: City, suburban, town, rural (categorical match)
- Grade Span: Elementary, middle, high, combined (categorical match)
- State: Same state preferred (configurable to national)
- Per-Pupil Expenditure: F-33 instructional spending per student
3.2 Distance Metric
We use a modified Gower distance that handles both continuous and categorical variables:
The 5 most similar schools are returned as peers. Users can adjust weights to prioritize certain matching criteria.
4. Funding Eligibility Rules
ThinkKits determines funding eligibility using the same criteria that federal agencies use for allocation formulas. We do not estimate eligibility — we apply the published rules.
4.1 Title I (ESEA Section I)
- Threshold: Schools where ≥40% of students come from low-income families (based on SAIPE district poverty data)
- Data source: Census SAIPE poverty estimates + CCD enrollment
- Allocation formula: Four formula grants (Basic, Concentration, Targeted, Education Finance Incentive) weighted by poverty count and state per-pupil expenditure
- Key rule: "Supplements, not supplants" — Title I funds must add to, not replace, existing spending
4.2 Title IV-A (Student Support and Academic Enrichment)
- Threshold: All LEAs receiving Title I are eligible
- Uses: Well-rounded education, safe/healthy schools, technology
- Key rule: At least 20% must go to "well-rounded education activities" if receiving ≥$30,000
4.3 IDEA Part B (Special Education)
- Threshold: All LEAs serving students with disabilities ages 3-21
- Data source: OSEP child count data + state allocation tables
- Key rule: Maintenance of effort — must spend at least as much on special ed as prior year
- Relevant for: Multi-sensory learning materials, assistive technology, IEP-aligned manipulatives
4.4 E-Rate (Universal Service)
- Eligibility: All schools and libraries can apply
- Discount rate: 20-90% based on FRL percentage and rural/urban status
- Categories: Category 1 (internet access, transport) and Category 2 (internal connections, managed Wi-Fi)
- Data source: USAC Open Data — commitment amounts, vendor SPINs, service types
5. Bright Spot Identification
A "Bright Spot" is a school that achieves better-than-expected outcomes given its demographic and funding profile. We identify them using a residual analysis approach:
- Build a regression model predicting performance (EDFacts proficiency) from demographics (FRL rate, enrollment, locale) and funding (per-pupil expenditure)
- Calculate the residual for each school: actual performance minus predicted performance
- Schools with residuals in the top 10% (outperforming expectations by the widest margin) are flagged as Bright Spots
This approach ensures Bright Spots aren't simply affluent schools with high scores — they are schools doing more with less or achieving outsized results given their context.
Bright Spot analysis requires EDFacts proficiency data, which is not available for all schools. Schools without assessment data are excluded from Bright Spot calculations but are still included in all other tools.
6. Update Cadence
| Data Source | Refresh Frequency | Typical Lag |
|---|---|---|
| NCES CCD | Annual | ~12-18 months (SY 2024-25 data available fall 2025) |
| E-Rate (USAC) | Quarterly | ~1-2 months from filing |
| SAIPE Poverty | Annual | ~18 months |
| CRDC | Biennial | ~24 months |
| EDFacts | Annual | ~12-18 months |
| F-33 Finance | Annual | ~18-24 months |
| Grants.gov | Daily | Same day |
| USASpending | Monthly | ~30 days |
ThinkKits ingests new releases within 48 hours of publication. Derived metrics (Health Score, Bright Spots) are recalculated after each major data refresh.
7. Limitations & Caveats
- Data lag: Federal education data is inherently lagged. The most recent CCD data is typically 12-18 months old. Real-time enrollment and staffing changes are not reflected.
- Missing data: Not all schools report all fields. Small schools and new charter schools are most likely to have incomplete records.
- Aggregate only: All data is school- or district-level. We cannot and do not provide student-level data.
- State variability: Assessment data (EDFacts) uses state tests, which vary in rigor and cut scores. Cross-state performance comparisons should be interpreted cautiously.
- Health Score: The composite metric is an analytical tool, not a rating. Context matters — a low score does not mean a school is "bad," and a high score does not guarantee quality.
- Peer matching: Statistical similarity does not imply identical contexts. Peer comparisons should supplement, not replace, qualitative knowledge of schools.
- Funding estimates: Allocation data shows what was allocated to districts, not necessarily what reaches individual schools. Intra-district distribution varies.