{
  "version": 1,
  "scope": "New editorial pages only. Not an audit of legacy site claims.",
  "claims": {
    "editorial-navigation": {
      "basis": "decided",
      "claim": "Publication navigation, categories, dates and reading links for the editorial collection.",
      "source": "src/editorial/articles.json",
      "first_published": "2026-09-25",
      "archive_policy": "Article archive dates follow the underlying research or development; the newsletter uses its edition month. Retrospective dates are not first-publication dates."
    },
    "september-2026-evidence-and-explanations": {
      "basis": "cited",
      "claim": "Evidence, explanations and what changed at Ackren. A September research digest: an inspectable teaching example, the limits of replay, and recent work on testing AI explanations.",
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        },
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        },
        {
          "label": "Pham, Le and Luu, GRACE, preprint, 15 June 2026",
          "url": "https://arxiv.org/html/2606.16151v1",
          "basis": "cited"
        },
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        }
      ],
      "archive_date": "2026-09-25",
      "first_published": "2026-09-25",
      "date_basis": "Monthly newsletter edition"
    },
    "september-2026-evidence-and-explanations-1": {
      "basis": "cited",
      "claim": "From the Ackren record",
      "text": [
        "Our latest research notes follow the path from a taught statement to an answer. The local prototype supplies a concrete example with visible evidence, while the earlier evaluation shows why repeatability and correct meaning need separate tests. Both articles keep the limits of their evidence alongside the result.",
        "<a href=\"/news/teaching-with-visible-evidence\">Read the prototype update</a> · <a href=\"/news/repeatable-is-not-correct\">Read the evaluation analysis</a>"
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "september-2026-evidence-and-explanations-2": {
      "basis": "cited",
      "claim": "From the research literature",
      "text": [
        "We review a recent approach to challenging visual explanations, alongside work on textual reasoning and language-to-planning translation. The articles distinguish the researchers’ results from our proposed applications to Ackren.",
        "<a href=\"/news/test-the-evidence-in-ai-explanations\">Read the explanation review</a> · <a href=\"/news/verification-starts-before-the-solver\">Read the verification analysis</a>"
      ],
      "sources": [
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        },
        {
          "label": "Pham, Le and Luu, GRACE, preprint, 15 June 2026",
          "url": "https://arxiv.org/html/2606.16151v1",
          "basis": "cited"
        },
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "september-2026-evidence-and-explanations-3": {
      "basis": "decided",
      "claim": "What would make the next result stronger",
      "text": [
        "Our editorial priority is independent evidence: unfamiliar language supplied after the candidate is frozen, complete examples including failures, and human assessment of the interpretation and explanation. These are proposed next steps. This edition reports engineering progress and research questions, not a general-conversation breakthrough.",
        "This is the web edition. Future editions can be followed through the <a href=\"/news/feed.xml\">RSS feed</a>."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    },
    "teaching-with-visible-evidence": {
      "basis": "cited",
      "claim": "Teaching a conversational engine, with the evidence visible. A recorded local example shows Ackren deriving an answer from a statement and a rule. The prototype remains bounded and has not completed release qualification.",
      "sources": [
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "archive_date": "2026-09-25",
      "first_published": "2026-09-25",
      "date_basis": "Publication date of the prototype update"
    },
    "teaching-with-visible-evidence-1": {
      "basis": "measured",
      "claim": "A small exchange with a visible derivation",
      "text": [
        "In a recorded local browser example, the user tells Ackren: “Mara is a composer.” The next statement supplies a rule: “Every composer is a person.” Asked “Is Mara a person?”, the prototype replies: “Mara is a person.” Its displayed explanation links the answer to the supplied statement and rule.",
        "The browser record also showed the selected proof, inference checks and word-generation trace. This is synthetic teaching data from a development example. It demonstrates the recorded exchange, not independent conversational competence or a verified fact about a real person."
      ],
      "sources": [
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "teaching-with-visible-evidence-2": {
      "basis": "cited",
      "claim": "What changed in the prototype",
      "text": [
        "The latest development record describes a non-neural local path that composes replies and explanations from explicit meaning and language records. It supports bounded teaching, inference and correction. The record also describes tentative relational proposals and review of changed interpretations. Those capabilities remain constrained by the implemented language and reasoning rules.",
        "A subsequent targeted repair addressed greeting-prefixed identity questions and remembered names written in lowercase. This is a concrete usability improvement, but a fix for a selected question form is not evidence of unrestricted language coverage."
      ],
      "sources": [
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "teaching-with-visible-evidence-3": {
      "basis": "counted",
      "claim": "The tests that still matter",
      "text": [
        "The prototype’s initial structural progress set recorded 46 passing turns out of 47. A broader exposed regression recorded 28 out of 39. These sets differ in construction, scope and exposure; the fractions must not be compared as if they measured an improvement or decline on one benchmark.",
        "The development record leaves relevant follow-up questions, broad-topic conversation and independent release qualification open. Passing engineering checks is valuable evidence about specified behaviour. It does not replace independent human evaluation."
      ],
      "sources": [
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "teaching-with-visible-evidence-4": {
      "basis": "decided",
      "claim": "Why this is worth following",
      "text": [
        "Our interest is in whether a person can teach, inspect and correct a system without losing the trail from a statement to a later answer. The example makes that question concrete. The next useful demonstration would include a correction, an ambiguous statement and an unsupported request, with the complete record available for each.",
        "We regard this as a research update with a testable direction. We are not presenting symbolic conversation as a new invention, claiming an advantage over language models, or treating a local candidate as a released service."
      ],
      "sources": [
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    },
    "repeatable-is-not-correct": {
      "basis": "cited",
      "claim": "Repeatable is not the same as correct. Ackren’s recorded evaluation separates repeatable execution from correct interpretation. That distinction matters when assessing an AI explanation.",
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "archive_date": "2026-09-24",
      "first_published": "2026-09-25",
      "date_basis": "Development record date"
    },
    "repeatable-is-not-correct-1": {
      "basis": "cited",
      "claim": "The result worth examining",
      "text": [
        "An earlier Ackren candidate replayed all 716 recorded turns identically, yet qualified on only 94 of 120 scored episodes. It passed 14 of 40 meaning-contrast episodes. These are counts from an exposed, agent-authored regression set, not estimates of performance with people. The candidate failed its overall and contrast qualification thresholds.",
        "This is a useful result to publish because the successful replay and the unsuccessful language evaluation concern different properties. A system can follow the same steps every time and still misunderstand the sentence it was given. Repeatability makes a failure easier to reproduce; it does not make the answer true."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "repeatable-is-not-correct-2": {
      "basis": "cited",
      "claim": "What the checks established",
      "text": [
        "The recorded run contains 476 scored checkpoints and 240 diagnostic turns. All 716 turns passed the structural explanation audit. That audit checked the shape of the explanation and its public record and source references. It did not establish that every explanatory sentence was faithful, that the English interpretation was correct, or that a person would find the conversation useful.",
        "The contrast cases are especially revealing. They change meaning while keeping much of the wording similar. The candidate qualified on 40 of 40 paraphrase episodes, but only 14 of 40 contrast episodes. A parser that tolerates alternate wording must also preserve consequential differences between statements. Success on the former does not establish the latter."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "repeatable-is-not-correct-3": {
      "basis": "cited",
      "claim": "Why the denominator matters",
      "text": [
        "The fixtures had been exposed during development, and the evaluator itself required repairs. The final figures describe regression performance on that material. They are not a fresh generalisation test. No human participants took part. Earlier partial evaluator scores must not be presented as a clean before-and-after improvement curve.",
        "These results belong to the frozen candidate identified in the evidence note, not to every Ackren build. The newer compositional prototype has different evidence and outstanding qualification work. Combining their strongest numbers would create a result that no single candidate achieved."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "repeatable-is-not-correct-4": {
      "basis": "decided",
      "claim": "The question for an AI demonstration",
      "text": [
        "Our editorial conclusion is that a convincing demonstration should expose the interpretation as well as the answer. Ask what the system understood, which premises it used, whether a changed premise changes the result appropriately, and how it responds when the wording is outside its coverage.",
        "A stronger follow-up study would freeze the implementation and evaluator before independent people supply unfamiliar examples. It would assess intended meaning and explanation usefulness separately from replay. That is proposed work. The existing record supports an inspectable engineering case study, not a claim that Ackren has solved general conversation."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    },
    "test-the-evidence-in-ai-explanations": {
      "basis": "cited",
      "claim": "To test an AI explanation, change the evidence it cites. A September preprint tests visual explanations through controlled image changes. The wider lesson is to test explanations against evidence, with clear limits on what a passing test proves.",
      "sources": [
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        },
        {
          "label": "Pham, Le and Luu, GRACE, preprint, 15 June 2026",
          "url": "https://arxiv.org/html/2606.16151v1",
          "basis": "cited"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "archive_date": "2026-09-16",
      "first_published": "2026-09-25",
      "date_basis": "Retrospective archive date matching the reviewed paper version"
    },
    "test-the-evidence-in-ai-explanations-1": {
      "basis": "cited",
      "claim": "The September research",
      "text": [
        "EDCT-Bench, a preprint posted on 16 September 2026 by Sihao Ding and colleagues, edits visual evidence cited in model explanations and tests the resulting responses. The authors report inconsistencies across the evaluated models.",
        "An answer can remain valid if other evidence supports it, but its explanation must account for the edit. Passing does not prove internal causal faithfulness. The preliminary training analysis does not establish improved held-out faithfulness."
      ],
      "sources": [
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "test-the-evidence-in-ai-explanations-2": {
      "basis": "cited",
      "claim": "A related test for textual reasoning",
      "text": [
        "The June GRACE preprint by Hoang Pham, Dong Le and Anh Tuan Luu examines individual reasoning steps against supplied context. It distinguishes deduction errors from failures to remain grounded in the source. The work reinforces an important evaluation distinction: a correct final answer can coexist with unsupported intermediate reasoning.",
        "These studies ask different questions. GRACE assesses textual reasoning against context; EDCT intervenes on visual evidence cited by a model. Their results should not be combined into a single accuracy number or used as direct evidence about Ackren."
      ],
      "sources": [
        {
          "label": "Pham, Le and Luu, GRACE, preprint, 15 June 2026",
          "url": "https://arxiv.org/html/2606.16151v1",
          "basis": "cited"
        },
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "test-the-evidence-in-ai-explanations-3": {
      "basis": "decided",
      "claim": "Our reading: make explanations answerable to evidence",
      "text": [
        "For a developer, the useful question is practical: what observation would reveal that an explanation is wrong? Reading a plausible paragraph is a weak test if the paragraph cannot be challenged. Changing a cited premise, keeping unrelated evidence fixed and checking the resulting answer gives a more specific test.",
        "For Ackren, an analogous experiment would revise a supplied fact or withdraw a rule, then inspect the answer and its dependencies. That is our proposed application of this research, not a result from either paper. An execution log makes such an experiment inspectable, but does not excuse errors in interpretation or in the log itself."
      ],
      "sources": [
        {
          "label": "Ding and colleagues, EDCT-Bench, preprint, 16 September 2026",
          "url": "https://arxiv.org/html/2609.17953v1",
          "basis": "cited"
        },
        {
          "label": "Pham, Le and Luu, GRACE, preprint, 15 June 2026",
          "url": "https://arxiv.org/html/2606.16151v1",
          "basis": "cited"
        },
        {
          "label": "Recorded prototype example and development checks",
          "url": "/news/evidence/ackren-prototype.json",
          "basis": "measured"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    },
    "verification-starts-before-the-solver": {
      "basis": "cited",
      "claim": "Verification starts before the solver. Research on language-to-planning systems highlights a difficult boundary: a formally checked result depends on what the system translated from the user’s request.",
      "sources": [
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        },
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        }
      ],
      "archive_date": "2026-07-04",
      "first_published": "2026-09-25",
      "date_basis": "Retrospective archive date matching the reviewed paper version"
    },
    "verification-starts-before-the-solver-1": {
      "basis": "cited",
      "claim": "The research question",
      "text": [
        "Jiang and colleagues’ Language Models as Higher-Order Planning Formalizers, revised in July 2026, studies a bottleneck in translating concise natural-language descriptions into planning representations. A short description can imply a much larger formal problem, making direct translation difficult.",
        "The authors propose generating a higher-level program that expands into the planning representation. They report improved performance on their constructed complex planning problems. This is evidence for a particular formalisation approach and evaluation setting, not a demonstration that arbitrary requests can be translated correctly."
      ],
      "sources": [
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        }
      ],
      "qualification": "Scope and limitations stated in the article."
    },
    "verification-starts-before-the-solver-2": {
      "basis": "decided",
      "claim": "The boundary a proof cannot skip",
      "text": [
        "Our reading is that verification needs to begin with the representation of the task. A solver can produce a valid result for the formal problem it receives while the overall system fails the request that a person intended. Testing the solver and testing the translation are distinct jobs.",
        "Consider an illustrative instruction: every submitted item needs review. A translation that substitutes “some items” changes the obligation before any solver runs. A later proof about that altered formal statement cannot establish that the original instruction was satisfied. This is an explanatory example, not a reported experiment from the paper."
      ],
      "sources": [
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    },
    "verification-starts-before-the-solver-3": {
      "basis": "decided",
      "claim": "How that applies to Ackren",
      "text": [
        "Ackren’s earlier evaluation provides a local reason to keep those checks separate. Its replay and structural explanation checks passed across all recorded turns, while the candidate failed the meaning-contrast qualification threshold. A deterministic front end also needs semantic evaluation.",
        "We would therefore ask for separate evidence about the interpreted request, the resulting inference and the explanation shown to the user. This is our editorial recommendation. There is no comparative experiment here showing that Ackren outperforms the planning systems in the paper."
      ],
      "sources": [
        {
          "label": "Recorded Ackren evaluation: counts, scope and source hashes",
          "url": "/news/evidence/ackren-evaluation.json",
          "basis": "counted"
        },
        {
          "label": "Jiang and colleagues, Language Models as Higher-Order Planning Formalizers, revised 4 July 2026",
          "url": "https://arxiv.org/html/2603.23844v2",
          "basis": "cited"
        }
      ],
      "qualification": "Editorial conclusions and proposed tests are not measured results."
    }
  }
}
