Charter and work programme, version 1.0
A binding description of what we will do
The manifesto says what we believe. This document says what we will do, what we will not do, who decides, and under which conditions we dissolve ourselves.
Grounds for founding
The Foundation exists to fill a gap: the measurement of AI capability, and the translation of those measurements into public argument, is today largely in the hands of the institutions that produce what is being measured.
Frontier labs evaluate their own models against their own benchmarks, write their own safety frameworks, and audit their own compliance. This is not evidence of bad faith — it is the consequence of missing capacity. The number of organisations doing independent measurement fits on one hand, and none of them works in Turkish.
The second reason is more uncomfortable. In a 2026 interview study, 17 of the 25 leading researchers interviewed expect systems with advanced R&D capability to be increasingly held inside companies and never shown publicly. If that is right, external measurement will be structurally behind — and will also be the only remaining instrument of scrutiny.
Purpose and limits
We do
- 2.1Independent capability measurementWe measure the capabilities of publicly available models using published, reproducible protocols.
- 2.2Primary-source translationWe translate frontier labs' safety frameworks, foundational academic work and policy documents into Turkish and keep them in a permanent archive.
- 2.3Forecast registryWe record dated, probabilistic forecasts and score their outcomes publicly.
- 2.4Document auditWe track changes, inconsistencies and version history in companies' own commitment documents, and report them.
- 2.5Policy contributionOn request we provide technical advice to public bodies — free of charge and published.
We do not
- 2.6We do not build modelsWe do not train frontier models. Producing what you measure devalues the measurement.
- 2.7We do not give single datesWe do not publish point forecasts of the form “AGI arrives on date X.” Only probability distributions with resolution criteria.
- 2.8We do not consult privatelyWe produce no technical advice that will not be published. No exceptions.
- 2.9We do not give investment adviceOur measurements are not designed or licensed to be converted into investment decisions.
Research programme
Four lines of work. Each has its own output, its own publication rhythm, and its own failure criterion.
| Line | What it does | Output | Rhythm |
|---|---|---|---|
| A · Measurement | Task horizon, compositional generalisation and agent reliability on public models. Independent reproduction of the METR and ARC protocols. | Raw data + method + code | every 6 months |
| B · Translation & archive | Turkish editions of safety frameworks, foundational papers and policy documents, with version history. | Archive + version diffs | continuous |
| C · Forecast registry | Dated, probabilistic forecasts with written resolution criteria; annual retrospective scoring (Brier). | Registry + score report | annual |
| D · Document audit | Monitoring company commitment documents; detecting version changes and weakenings. | Change bulletin | quarterly |
Governance
- 4.1BoardFive members. At least two from non-technical fields (law, economics, public policy). Three-year terms, maximum two terms.
- 4.2Scientific advisory panelAt least seven people. At least two are drawn from researchers who openly reject the Foundation's core thesis. This is a binding rule, not a courtesy.
- 4.3Editorial independenceThe Board may not alter the content of a research output. It may only delay publication, and the reason for the delay is published alongside the output itself.
- 4.4Decision thresholdsOrdinary decisions by simple majority. Charter amendments, dissolution and changes to the income policy require a three-quarters majority.
- 4.5Dissenting opinionsOn any Board decision, members in the minority may attach written reasons; those dissents are published with the decision.
Transparency and conflicts of interest
- 5.1Income disclosureAll sources of income are published quarterly with amounts and conditions. Anonymous donations are not accepted.
- 5.2Lab ceilingTotal donations from frontier AI labs and their direct investors may not exceed 25% of annual income. Any excess is returned.
- 5.3Personal disclosureBoard members and researchers declare shares, options and advisory relationships with AI companies in a public register.
- 5.4Measurement embargoIf a company's model is measured in a period in which the Foundation received a donation from that company, the result is published together with, and with reference to, that disclosure.
- 5.5Raw dataThe raw data of every measurement is published under an open licence at the same time as the processed result.
Publication standard
No publication of the Foundation may go out without meeting the five rules below.
- 6.1Strongest objectionEvery publication states the strongest available objection to its own thesis, in a form its advocate would accept.
- 6.2Evidence gradingEvery numerical claim is tagged with a robustness grade: strong (multi-source, measured) · moderate (single source or methodological caveat) · weak (contradictory or framing-sensitive).
- 6.3Error logClaims that turn out to be wrong are not deleted. They stay published with the date and reason for correction.
- 6.4Method alongsideResult, method and raw data are published together. A result without a method may not be published.
- 6.5The limit of the measureBeside every measurement we state what that measurement does not show.
Measurement methodology annex
What we measure, what we measure it with, and which trap we know about.
| Indicator | What it measures | Known trap |
|---|---|---|
| Task time horizon — METR protocol | The human-expert duration of the task a model completes with 50% success. | Can go to infinity without taking other measures of intelligence with it (Ord). Highly sensitive to the composition of the task pool. |
| Compositional generalisation — ARC-AGI family | Efficiency at acquiring new rule structures from few examples. | Not comparable without stating the compute budget: on the same benchmark a constrained competition scores 24%, an unconstrained frontier system 68.8%. |
| Generation time — Ord's proposal | The duration of one turn of the R&D feedback loop. | No lab currently reports it. The Foundation campaigns for this (see 9.2). |
| Expert survey — Grace et al. type | Researchers' probability distributions. | Extremely framing-sensitive: the median differs by more than 69 years between “high-level machine intelligence” and “full automation of labour.” |
| Crowd forecast — Metaculus type | The aggregate distribution of scored forecasters. | Question-dependent; produces a 5-year gap between “general” and “weakly general.” |
| Scaling constraints — Epoch type | Modelling of power, chip, data and latency limits. | Measures feasibility, not intent. The gap between “possible” and “will happen” is economic and political. |
Three-year roadmap
2026 Q4
- First public release of the archive: 766 studies, 17 themes, full-text sources
- Opening of the forecast registry and its public announcement
- Formation of the scientific advisory panel — including two members who reject our thesis
2027
- Line A: independent reproduction of the METR protocol on public models
- Line B: full Turkish translation of three safety frameworks, with version diffs
- Line D: first annual document-audit bulletin
- First annual Brier score report (forecasts 01, 02, 04, 05, 08 resolve)
2028
- Line A: reproduction of the ARC protocol with declared compute budgets
- Public advocacy campaign for generation-time reporting
- First technical advice to a public body — free of charge and published
- Second Brier report (forecasts 03, 06, 07, 09 resolve)
2029
- Three-year self-assessment: resolution of forecast 10
- Review of the charter; application of clause 10.2 if required
Dissolution and restructuring conditions
A foundation that writes no condition for dissolving itself has made existing its purpose. The conditions below are binding.
- 10.1Publication silenceIf the Foundation publishes no research output for 24 months, it dissolves automatically. Existing cannot substitute for working.
- 10.2Method failureIf the forecast registry's Brier score fails to beat the naive baseline over a three-year period, the result is published and the Board restructures the method. If it fails again over a second three-year period, the Foundation dissolves.
- 10.3Loss of independenceIf the 25% ceiling in clause 5.2 is exceeded in two consecutive years, the Foundation ceases operations and publishes the reason.
- 10.4Purpose achievedIf independent, public, Turkish-accessible capability measurement is taken up by a permanent public institution, the Foundation transfers its archive to that body and dissolves. This is not a defeat; it is the success scenario.
- 10.5Purpose voidedPer article X of the manifesto: if the trajectory of this technology stops treating the human as the end, everything the Foundation defends becomes void, and the Board is obliged to establish that.