OpenAI Models: Behavior, Reliability, and Accountability
Status · Active research
This program examines OpenAI models and products through preserved conversations, exported records, technical logs, reproducible failures, and public product behavior. It separates what a model says about itself from what the record can actually establish.
Central question
When OpenAI systems claim progress, capability, correction, or accountability, what has the available evidence actually earned?
Current evidence
Featured working paper · August 27, 2026 · Not peer reviewed
Task completion without consequence
Evidence: A claim matrix compares the January 2025 warning and May 2026 article against the later incident record, preserving direct matches, partial matches, misses, contradictions, and untested claims.
Boundary: This is a retrospective comparison, not a preregistered forecast evaluation; it does not claim prediction of the exact exploit chain or establish independent replication.
OpenAI Investigation · Article 1
Why I Am Done With OpenAI — and How Its Support System Failed Me
Evidence: A catastrophic long-chat failure, a 71 GB quarantined conversation, local diagnosis and recovery, preserved support correspondence, and the absence of a meaningful human escalation path.
Boundary: The public optimization announcement does not prove that this individual report caused the update.
OpenAI Investigation · Article 2
GPT-5.5 Kept Me Working: False Progress, Capability Confusion, and the Engagement Loop
Evidence: An official ChatGPT export containing exact GPT-5.5-labeled messages, premature success declarations, repeated failed repair loops, and acknowledgments that did not produce durable behavioral correction.
Boundary: Model behavior can document an engagement-seeking pattern; it does not by itself prove OpenAI’s corporate intent.
Forensic working paper
Compaction-associated image payload amplification in a 71 GiB Codex rollout
Evidence: privacy-safe full-file byte attribution, SHA-256 duplicate grouping, known-answer fixture validation, and recurrence across compaction records.
Boundary: The storage mechanism is established for this artifact; the precise runtime path to SIGKILL remains open.
Incident analysis · article
The swarm was not magic. It was architecture.
Evidence: OpenAI, Hugging Face, and METR/Redwood records support several load-bearing mechanisms and meaningful overlap with the public May warning.
Boundary: The analysis earns architectural legibility, not prospective predictability or exclusive priority.
Protocol draft
Agentic boundary expansion: experimental protocol
The controlled study still needed to isolate task pressure, safe exits, shared state, peer authority, credentials, infrastructure, and correction.
Method
Preserve the primary record → identify the exact claim → separate observation from inference → search for contradiction and counterevidence → state what remains unproved → revise when new evidence arrives.
Current limits
This is an owner-built research record, not an independent audit of OpenAI’s internal systems. Private evidence remains private when publication would expose credentials, unrelated people, medical material, or local infrastructure.