OK: Found an XML parser.
OK: Support for GZIP encoding.
OK: Support for character munging.

Example Output

Channel: North Carolina Chronicle

RSS URL:

Parsed Results (var_dump'ed)

object(MagpieRSS)#4 (22) {
  ["parser"]=>
  resource(10) of type (Unknown)
  ["current_item"]=>
  array(0) {
  }
  ["items"]=>
  array(10) {
    [0]=>
    array(11) {
      ["title"]=>
      string(72) "Charlotte CATS rider hit and killed by bus she had just left, police say"
      ["link"]=>
      string(68) "https://nocarolinachronicle.com/charlotte-cats-rider-hit-killed-bus/"
      ["dc"]=>
      array(1) {
        ["creator"]=>
        string(10) "Bill Moran"
      }
      ["pubdate"]=>
      string(31) "Sat, 03 Oct 2026 19:50:11 +0000"
      ["category"]=>
      string(64) "NewsCharlotte Area TransitNorth Graham StreetPedestrian Accident"
      ["guid"]=>
      string(68) "https://nocarolinachronicle.com/charlotte-cats-rider-hit-killed-bus/"
      ["description"]=>
      string(1013) "
Charlotte CATS rider struck
A 49-year-old woman was struck and killed by a Charlotte Area Transit System bus she had just exited on North Graham Street, police report." ["content"]=> array(1) { ["encoded"]=> string(3521) "
Charlotte CATS rider struck

A 49-year-old woman was struck and killed by a Charlotte Area Transit System bus she had just exited Friday afternoon in the 1300 block of North Graham Street in Charlotte, police said. According to the Charlotte-Mecklenburg Police Department, the woman walked back toward the bus as it was moving away from the stop and was hit in the roadway; the investigation is ongoing.

The victim was identified by the Charlotte-Mecklenburg Police Department as 49-year-old Ebony Wingate. Officers arrived at the scene in the 1300 block of North Graham Street, near the intersection with Dalton Avenue, at approximately 1:10 p.m. on Friday, October 2, 2026. According to police,

Wingate was found unresponsive in the roadway and was pronounced dead at the scene by emergency medical personnel.

Preliminary information from CMPD indicates that Wingate had just exited a Charlotte Area Transit System bus at a designated stop when she walked back toward the bus as it was pulling away. The bus had begun moving as Wingate entered the street, at which point she was struck by the vehicle. The department described this account as preliminary and said the investigation is ongoing.

No other injuries or victims were reported in connection with the incident. CMPD’s Major Crash Investigation Unit has taken over the investigation, categorizing the case as a fatal crash involving a transit bus and pedestrian. As of October 3, police had not announced whether any charges or citations had been issued to the bus driver.

The crash location is situated north of Uptown Charlotte, a detail confirmed by local news sources. The Charlotte Observer reported the incident on the morning of October 3, with an update later that afternoon. WBTV also covered the crash, citing CMPD as the source for the timeline and details.

Authorities have asked anyone with information or who may have witnessed the collision to contact Detective Justin Kupfer at 704-432-2169, extension 1. Anonymous tips can be submitted to Charlotte Crime Stoppers at 704-334-1600. No additional witness statements have been publicly released.

Further details about Wingate’s background or residence have not been provided by the police or in public reports. The investigation remains active as officials continue to gather information and determine the full circumstances surrounding the fatal collision.

.

" } ["summary"]=> string(1013) "
Charlotte CATS rider struck
A 49-year-old woman was struck and killed by a Charlotte Area Transit System bus she had just exited on North Graham Street, police report." ["atom_content"]=> string(3521) "
Charlotte CATS rider struck

A 49-year-old woman was struck and killed by a Charlotte Area Transit System bus she had just exited Friday afternoon in the 1300 block of North Graham Street in Charlotte, police said. According to the Charlotte-Mecklenburg Police Department, the woman walked back toward the bus as it was moving away from the stop and was hit in the roadway; the investigation is ongoing.

The victim was identified by the Charlotte-Mecklenburg Police Department as 49-year-old Ebony Wingate. Officers arrived at the scene in the 1300 block of North Graham Street, near the intersection with Dalton Avenue, at approximately 1:10 p.m. on Friday, October 2, 2026. According to police,

Wingate was found unresponsive in the roadway and was pronounced dead at the scene by emergency medical personnel.

Preliminary information from CMPD indicates that Wingate had just exited a Charlotte Area Transit System bus at a designated stop when she walked back toward the bus as it was pulling away. The bus had begun moving as Wingate entered the street, at which point she was struck by the vehicle. The department described this account as preliminary and said the investigation is ongoing.

No other injuries or victims were reported in connection with the incident. CMPD’s Major Crash Investigation Unit has taken over the investigation, categorizing the case as a fatal crash involving a transit bus and pedestrian. As of October 3, police had not announced whether any charges or citations had been issued to the bus driver.

The crash location is situated north of Uptown Charlotte, a detail confirmed by local news sources. The Charlotte Observer reported the incident on the morning of October 3, with an update later that afternoon. WBTV also covered the crash, citing CMPD as the source for the timeline and details.

Authorities have asked anyone with information or who may have witnessed the collision to contact Detective Justin Kupfer at 704-432-2169, extension 1. Anonymous tips can be submitted to Charlotte Crime Stoppers at 704-334-1600. No additional witness statements have been publicly released.

Further details about Wingate’s background or residence have not been provided by the police or in public reports. The investigation remains active as officials continue to gather information and determine the full circumstances surrounding the fatal collision.

.

" ["date_timestamp"]=> int(1791057011) } [1]=> array(11) { ["title"]=> string(77) "Two deputies shot while confronting man firing at neighbors in North Carolina" ["link"]=> string(81) "https://nocarolinachronicle.com/two-deputies-shot-confronting-man-north-carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Sat, 03 Oct 2026 19:49:36 +0000" ["category"]=> string(39) "NewsDeep RunLenoir CountyNorth Carolina" ["guid"]=> string(81) "https://nocarolinachronicle.com/two-deputies-shot-confronting-man-north-carolina/" ["description"]=> string(1058) "
Deputies shot during confrontation
Two Lenoir County deputies were shot while responding to a man firing at neighbors in Deep Run, NC; the suspect was later fatally shot." ["content"]=> array(1) { ["encoded"]=> string(4528) "
Deputies shot during confrontation

Two Lenoir County deputies were shot Friday morning while responding to reports of a man firing at neighbors in the 700 block of Jonestown Road near Deep Run, North Carolina. According to Sheriff Jackie Rogers, the suspect opened fire on officers upon their arrival and later barricaded himself inside a residence before being fatally shot by deputies.

The two deputies wounded in the shooting were identified as Capt. Jovanni Villagra and Sgt. Cole Davis, according to Lenoir County Sheriff Jackie Rogers. Villagra was shot in the stomach during the initial exchange of gunfire after the suspect, later identified as 58-year-old Cecil James Langston of Deep Run, emerged from the residence and fired at officers upon their arrival. Langston then retreated inside the home, where he fired again, striking Sgt. Davis before officers entered the residence.

Both deputies were transported to ECU Health Medical Center in Greenville, where officials reported their injuries were non-life-threatening and that they remained in stable condition.

The incident began when deputies responded to reports of a man firing at people passing by in the 700 block of Jonestown Road near Deep Run on Friday, October 2, 2026. The call came in between approximately 11 a.m. and 11:20 a.m., according to Sheriff Rogers. Upon arrival, Langston opened fire on the deputies, prompting them to return fire and leading to a standoff at the residence. Additional law enforcement personnel from local and state agencies arrived to assist, and armored vehicles were used to breach the front door of the home.

Authorities deployed a drone to locate Langston inside the residence, where he was found in a bedroom. After officers entered, Langston was pronounced dead at the scene, according to the Lenoir County Sheriff’s Office and local reports. Sheriff Rogers confirmed the suspect’s death and stated that there was no continuing threat to the public following the conclusion of the incident.

Following the shooting, the North Carolina State Bureau of Investigation (SBI) took over the investigation, classifying the case as an officer-involved shooting due to multiple law enforcement officers discharging their weapons. The deputies involved were placed on administrative duty during the inquiry, which focuses on the exchanges of gunfire between Langston and responding officers. As of October 3, 2026, the SBI reported that the deputies remained alive and in stable condition, while the suspect was deceased and the scene was secured.

Sheriff Rogers emphasized that the immediate concern was the safety and well-being of Capt. Villagra and Sgt. Davis. He also advised the public to rely on official law enforcement updates as the investigation continued. Authorities confirmed that no ongoing danger to residents remained after the incident was contained.

The shooting occurred in the Pink Hill area of southern Lenoir County, near Deep Run, a rural community where such incidents are rare. The use of armored vehicles and drone technology highlighted the coordinated response by multiple agencies to ensure the safety of officers and the public. The SBI’s ongoing investigation will determine the full circumstances surrounding the shooting and the actions of all parties involved.

.

" } ["summary"]=> string(1058) "
Deputies shot during confrontation
Two Lenoir County deputies were shot while responding to a man firing at neighbors in Deep Run, NC; the suspect was later fatally shot." ["atom_content"]=> string(4528) "
Deputies shot during confrontation

Two Lenoir County deputies were shot Friday morning while responding to reports of a man firing at neighbors in the 700 block of Jonestown Road near Deep Run, North Carolina. According to Sheriff Jackie Rogers, the suspect opened fire on officers upon their arrival and later barricaded himself inside a residence before being fatally shot by deputies.

The two deputies wounded in the shooting were identified as Capt. Jovanni Villagra and Sgt. Cole Davis, according to Lenoir County Sheriff Jackie Rogers. Villagra was shot in the stomach during the initial exchange of gunfire after the suspect, later identified as 58-year-old Cecil James Langston of Deep Run, emerged from the residence and fired at officers upon their arrival. Langston then retreated inside the home, where he fired again, striking Sgt. Davis before officers entered the residence.

Both deputies were transported to ECU Health Medical Center in Greenville, where officials reported their injuries were non-life-threatening and that they remained in stable condition.

The incident began when deputies responded to reports of a man firing at people passing by in the 700 block of Jonestown Road near Deep Run on Friday, October 2, 2026. The call came in between approximately 11 a.m. and 11:20 a.m., according to Sheriff Rogers. Upon arrival, Langston opened fire on the deputies, prompting them to return fire and leading to a standoff at the residence. Additional law enforcement personnel from local and state agencies arrived to assist, and armored vehicles were used to breach the front door of the home.

Authorities deployed a drone to locate Langston inside the residence, where he was found in a bedroom. After officers entered, Langston was pronounced dead at the scene, according to the Lenoir County Sheriff’s Office and local reports. Sheriff Rogers confirmed the suspect’s death and stated that there was no continuing threat to the public following the conclusion of the incident.

Following the shooting, the North Carolina State Bureau of Investigation (SBI) took over the investigation, classifying the case as an officer-involved shooting due to multiple law enforcement officers discharging their weapons. The deputies involved were placed on administrative duty during the inquiry, which focuses on the exchanges of gunfire between Langston and responding officers. As of October 3, 2026, the SBI reported that the deputies remained alive and in stable condition, while the suspect was deceased and the scene was secured.

Sheriff Rogers emphasized that the immediate concern was the safety and well-being of Capt. Villagra and Sgt. Davis. He also advised the public to rely on official law enforcement updates as the investigation continued. Authorities confirmed that no ongoing danger to residents remained after the incident was contained.

The shooting occurred in the Pink Hill area of southern Lenoir County, near Deep Run, a rural community where such incidents are rare. The use of armored vehicles and drone technology highlighted the coordinated response by multiple agencies to ensure the safety of officers and the public. The SBI’s ongoing investigation will determine the full circumstances surrounding the shooting and the actions of all parties involved.

.

" ["date_timestamp"]=> int(1791056976) } [2]=> array(11) { ["title"]=> string(67) "The One Prompting Error You’re Making, and the Research Behind It" ["link"]=> string(110) "https://nocarolinachronicle.com/pinkelephant-negative-prompting-dont-instructions-llm-research-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Fri, 02 Oct 2026 15:06:49 +0000" ["category"]=> string(151) "NewsAI agentsChatGPTClaudeironic reboundLlama 3LLMlong contextlost in the middlenegative promptingpink elephant problemprompt engineeringsystem prompts" ["guid"]=> string(110) "https://nocarolinachronicle.com/pinkelephant-negative-prompting-dont-instructions-llm-research-north_carolina/" ["description"]=> string(419) "The most common prompting mistake is telling an AI model what not to do. Research on negation, the Pink Elephant problem and ironic rebound shows naming a forbidden topic can make it more likely, with 15 to 20 middle-layer attention heads in Llama 3 driving the effect, and long contexts weaken early rules further. The fix: say what you want, place rules at the start or end, and enforce hard limits outside the model." ["content"]=> array(1) { ["encoded"]=> string(5781) "

By Bill Moran

The one prompting error almost everybody makes is telling the model what not to do. Don’t be verbose. Don’t make up sources. Don’t mention the competitor. Seven research papers on negation, priming and long context explain why “don’t” is the weakest instruction you can give a language model, and why it gets weaker the longer a session runs.

To obey “don’t mention X,” a model has to represent X, and that representation can make X more likely to appear. Novel Cognition read the research behind it. The short version: tell the model where to go, not where not to look.

The full walkthrough — why a bare prohibition gives a model nothing to aim at, how models handle negation, the Pink Elephant problem, the attention heads behind ironic rebound, Lost in the Middle, and what to write instead. Six minutes.

The most common prompting mistake is telling an AI model what not to do. Research on negation, the Pink Elephant problem and ironic rebound shows naming a forbidden topic can make it more likely, with

What the research shows

Studies from 2023 found GPT-3-era models often ignored negation and failed to reason under it; the NeQA benchmark found negated questions did not improve straightforwardly with scale, and a 2025 study found larger models handle it better but not reliably. Castricato et al.’s 2024 “Pink Elephant” paper found that mentioning a topic in the prompt made models more likely to mention it even when told not to: baseline models brought it up “marginally more frequently, and certainly not less.” Their Direct Principle Feedback fine-tune fixed it, bringing a 13B Llama 2 to GPT-4-level performance on their test.

A 2025 preprint, “Don’t Think of the White Bear,” built ReboundBench, 5,000 negation prompts across nine open models, and found the forbidden word rebounds immediately and more strongly with longer or on-topic text in between. In Llama-3-8B-Instruct, 15 to 20 middle-layer attention heads accounted for 80% or more of the rebound. Details at Naming It Primes It and Inside the Model.

Why long sessions make it worse

Liu et al.’s “Lost in the Middle” found models use information at the start and end of long inputs best and the middle worst. A rule set once at the top of a long session drifts into that middle as tool outputs and documents pile up, and on-topic documents are exactly the condition that produced the strongest rebound. Combined, prohibitions are the rules most likely to break deep into a long session. That conclusion is an inference from converging studies: no paper has tested negation rebound at 100,000 tokens, and the key study is a preprint using models up to 20B parameters. See Why Long Sessions Make It Worse and What the Evidence Doesn’t Show.

Say where to go

State the behavior you want (“answer in two sentences,” not “don’t be verbose”). Put critical rules at the start or end of a prompt, and restate essential ones near the end. Avoid naming the forbidden thing when you can describe the allowed alternative. Enforce hard constraints outside the model with filters, regex checks, structured outputs or logit bias, and re-inject key rules in long sessions. Then test on your own model at your real context length.

The full file is at pinkelephant.novcog.us.com. Primary sources: Castricato et al., Mann et al., Liu et al., and Truong et al.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily Texas News
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(419) "The most common prompting mistake is telling an AI model what not to do. Research on negation, the Pink Elephant problem and ironic rebound shows naming a forbidden topic can make it more likely, with 15 to 20 middle-layer attention heads in Llama 3 driving the effect, and long contexts weaken early rules further. The fix: say what you want, place rules at the start or end, and enforce hard limits outside the model." ["atom_content"]=> string(5781) "

By Bill Moran

The one prompting error almost everybody makes is telling the model what not to do. Don’t be verbose. Don’t make up sources. Don’t mention the competitor. Seven research papers on negation, priming and long context explain why “don’t” is the weakest instruction you can give a language model, and why it gets weaker the longer a session runs.

To obey “don’t mention X,” a model has to represent X, and that representation can make X more likely to appear. Novel Cognition read the research behind it. The short version: tell the model where to go, not where not to look.

The full walkthrough — why a bare prohibition gives a model nothing to aim at, how models handle negation, the Pink Elephant problem, the attention heads behind ironic rebound, Lost in the Middle, and what to write instead. Six minutes.

The most common prompting mistake is telling an AI model what not to do. Research on negation, the Pink Elephant problem and ironic rebound shows naming a forbidden topic can make it more likely, with

What the research shows

Studies from 2023 found GPT-3-era models often ignored negation and failed to reason under it; the NeQA benchmark found negated questions did not improve straightforwardly with scale, and a 2025 study found larger models handle it better but not reliably. Castricato et al.’s 2024 “Pink Elephant” paper found that mentioning a topic in the prompt made models more likely to mention it even when told not to: baseline models brought it up “marginally more frequently, and certainly not less.” Their Direct Principle Feedback fine-tune fixed it, bringing a 13B Llama 2 to GPT-4-level performance on their test.

A 2025 preprint, “Don’t Think of the White Bear,” built ReboundBench, 5,000 negation prompts across nine open models, and found the forbidden word rebounds immediately and more strongly with longer or on-topic text in between. In Llama-3-8B-Instruct, 15 to 20 middle-layer attention heads accounted for 80% or more of the rebound. Details at Naming It Primes It and Inside the Model.

Why long sessions make it worse

Liu et al.’s “Lost in the Middle” found models use information at the start and end of long inputs best and the middle worst. A rule set once at the top of a long session drifts into that middle as tool outputs and documents pile up, and on-topic documents are exactly the condition that produced the strongest rebound. Combined, prohibitions are the rules most likely to break deep into a long session. That conclusion is an inference from converging studies: no paper has tested negation rebound at 100,000 tokens, and the key study is a preprint using models up to 20B parameters. See Why Long Sessions Make It Worse and What the Evidence Doesn’t Show.

Say where to go

State the behavior you want (“answer in two sentences,” not “don’t be verbose”). Put critical rules at the start or end of a prompt, and restate essential ones near the end. Avoid naming the forbidden thing when you can describe the allowed alternative. Enforce hard constraints outside the model with filters, regex checks, structured outputs or logit bias, and re-inject key rules in long sessions. Then test on your own model at your real context length.

The full file is at pinkelephant.novcog.us.com. Primary sources: Castricato et al., Mann et al., Liu et al., and Truong et al.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily Texas News
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790953609) } [3]=> array(11) { ["title"]=> string(85) "A DGX Spark Cluster Was Built for an SEO Agency. Here’s Why It’s the Wrong Answer" ["link"]=> string(115) "https://nocarolinachronicle.com/sparkoverbuild-dgx-spark-cluster-seo-agency-mac-studio-har-analysis-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Tue, 29 Sep 2026 16:29:18 +0000" ["category"]=> string(107) "NewsAEOAI searchAI visibilityApple UpgradeDGX Sparkgpt-oss-120bHAR filelocal LLMM3 UltraMac StudioNVIDIASEO" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53045" ["description"]=> string(536) "A post announced four NVIDIA DGX Spark systems running local LLMs 24/7 for SEO and AI-search research, including HAR-file analysis of AI search products. The HAR idea is good; the hardware case isn't made. Most of the listed jobs are a data pipeline, a HAR only shows what crossed the browser boundary, four 128 GB Sparks are four memory pools rather than one 512 GB machine, and Exo Labs found the Spark faster at prefill but an M3 Ultra Mac Studio faster at generation. Apple now leases Macs, letting an agency measure before it buys." ["content"]=> array(1) { ["encoded"]=> string(6741) "

By Bill Moran

A DGX Spark cluster was build for an SEO agency. Here’s why that is the wrong answer.

A widely shared post announced four NVIDIA DGX Spark systems in a home server room, running local LLMs around the clock for SEO and AI-search research: vendor-claim tracking, pricing watch, SERP labelling, a “citability” model, contradiction detection, and analysis of HAR files captured from AI search products. Novel Cognition read it closely. The research agenda is credible. The case for the hardware is not made: the post offers page counts, keyword counts and a roughly $100 monthly power estimate, but no throughput, accuracy or utilization figures.

What a HAR file actually sees

A HAR (HTTP Archive) file logs a browser’s network traffic: requests, responses, timing and, depending on the capture, streamed events. When an AI search product streams its queries or candidate sources to the browser, a HAR can capture them, which makes it genuinely useful evidence. But the post claims HARs show “every search turn,” “every page it fetched,” and every page “fetched and then declined to cite.” Each of those requires the product to label that event in browser-visible traffic. Whatever the provider does on its own servers and never sends the browser, a HAR cannot see.

The research pays when it answers a decision a client will act on, measured by accepted findings rather than traces processed, with repeated runs, controls and deterministic parsing before any model is involved. See What a HAR File Actually Sees and Does HAR Analysis Pay?.

A post announced four NVIDIA DGX Spark systems running local LLMs 24/7 for SEO and AI-search research, including HAR-file analysis of AI search products. The HAR idea is good; the hardware case isn’t

The full walkthrough — what a HAR file can and can’t show about AI search, whether HAR analysis pays, why the SEO jobs are mostly a data pipeline, model choice and the harness, four DGX Sparks against one Mac Studio, and leasing instead of buying. Seven minutes.

Six jobs, mostly a pipeline

Sixty thousand pages a month averages about 2,000 a day; fetching, snapshotting, diffing and extracting them is data engineering. A weekly program of 14,930 keywords is mostly a SERP-collection problem. An LLM is useful for the ambiguous remainder after code has done its work, and the post never says how large that remainder is. Its citability model is trained on pages marked “cited but not retrieved,” a label that depends on retrieval most outside observers cannot see. What decides quality is the harness: snapshots, a typed evidence store, a changed-material queue, a cheap-first model router, schema-checked output and evaluation sets. Details at Six Jobs, Not One and Model Choice and the Harness.

Four Sparks against one Mac Studio

NVIDIA lists each DGX Spark with 128 GB of unified memory and 273 GB/s of bandwidth, so four provide 512 GB across four separate systems rather than one machine. An M3 Ultra Mac Studio holds up to 256 GB in a single pool. Exo Labs measured the Spark prefilling prompts 3.8 times faster than the M3 Ultra and the Mac generating tokens 3.4 times faster; the right choice depends on the bottleneck, which the post never measured. Apple now leases Macs through Apple Upgrade on 24- or 36-month terms (a personal lease; businesses use Apple’s business financing), which lets an agency measure its real workload before committing roughly $18,800 to four Sparks at the reported price.

The full comparison is at Four Sparks vs One Mac Studio and Lease the Mac, Measure First.

The question the post never answers

How much more billable, accurate, timely research does the fourth Spark produce over the first Mac Studio, and what does that earn? Until someone can answer, four Sparks are defensible as research or as marketing, not as an operating necessity. The moat is the experiment, the evidence and the analyst.

The full file is at sparkoverbuild.novcog.us.com. Primary sources: NVIDIA’s DGX Spark hardware overview, Exo Labs’ Spark and Mac Studio test, Apple’s Apple Upgrade announcement, and the W3C HAR specification.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(536) "A post announced four NVIDIA DGX Spark systems running local LLMs 24/7 for SEO and AI-search research, including HAR-file analysis of AI search products. The HAR idea is good; the hardware case isn't made. Most of the listed jobs are a data pipeline, a HAR only shows what crossed the browser boundary, four 128 GB Sparks are four memory pools rather than one 512 GB machine, and Exo Labs found the Spark faster at prefill but an M3 Ultra Mac Studio faster at generation. Apple now leases Macs, letting an agency measure before it buys." ["atom_content"]=> string(6741) "

By Bill Moran

A DGX Spark cluster was build for an SEO agency. Here’s why that is the wrong answer.

A widely shared post announced four NVIDIA DGX Spark systems in a home server room, running local LLMs around the clock for SEO and AI-search research: vendor-claim tracking, pricing watch, SERP labelling, a “citability” model, contradiction detection, and analysis of HAR files captured from AI search products. Novel Cognition read it closely. The research agenda is credible. The case for the hardware is not made: the post offers page counts, keyword counts and a roughly $100 monthly power estimate, but no throughput, accuracy or utilization figures.

What a HAR file actually sees

A HAR (HTTP Archive) file logs a browser’s network traffic: requests, responses, timing and, depending on the capture, streamed events. When an AI search product streams its queries or candidate sources to the browser, a HAR can capture them, which makes it genuinely useful evidence. But the post claims HARs show “every search turn,” “every page it fetched,” and every page “fetched and then declined to cite.” Each of those requires the product to label that event in browser-visible traffic. Whatever the provider does on its own servers and never sends the browser, a HAR cannot see.

The research pays when it answers a decision a client will act on, measured by accepted findings rather than traces processed, with repeated runs, controls and deterministic parsing before any model is involved. See What a HAR File Actually Sees and Does HAR Analysis Pay?.

A post announced four NVIDIA DGX Spark systems running local LLMs 24/7 for SEO and AI-search research, including HAR-file analysis of AI search products. The HAR idea is good; the hardware case isn’t

The full walkthrough — what a HAR file can and can’t show about AI search, whether HAR analysis pays, why the SEO jobs are mostly a data pipeline, model choice and the harness, four DGX Sparks against one Mac Studio, and leasing instead of buying. Seven minutes.

Six jobs, mostly a pipeline

Sixty thousand pages a month averages about 2,000 a day; fetching, snapshotting, diffing and extracting them is data engineering. A weekly program of 14,930 keywords is mostly a SERP-collection problem. An LLM is useful for the ambiguous remainder after code has done its work, and the post never says how large that remainder is. Its citability model is trained on pages marked “cited but not retrieved,” a label that depends on retrieval most outside observers cannot see. What decides quality is the harness: snapshots, a typed evidence store, a changed-material queue, a cheap-first model router, schema-checked output and evaluation sets. Details at Six Jobs, Not One and Model Choice and the Harness.

Four Sparks against one Mac Studio

NVIDIA lists each DGX Spark with 128 GB of unified memory and 273 GB/s of bandwidth, so four provide 512 GB across four separate systems rather than one machine. An M3 Ultra Mac Studio holds up to 256 GB in a single pool. Exo Labs measured the Spark prefilling prompts 3.8 times faster than the M3 Ultra and the Mac generating tokens 3.4 times faster; the right choice depends on the bottleneck, which the post never measured. Apple now leases Macs through Apple Upgrade on 24- or 36-month terms (a personal lease; businesses use Apple’s business financing), which lets an agency measure its real workload before committing roughly $18,800 to four Sparks at the reported price.

The full comparison is at Four Sparks vs One Mac Studio and Lease the Mac, Measure First.

The question the post never answers

How much more billable, accurate, timely research does the fourth Spark produce over the first Mac Studio, and what does that earn? Until someone can answer, four Sparks are defensible as research or as marketing, not as an operating necessity. The moat is the experiment, the evidence and the analyst.

The full file is at sparkoverbuild.novcog.us.com. Primary sources: NVIDIA’s DGX Spark hardware overview, Exo Labs’ Spark and Mac Studio test, Apple’s Apple Upgrade announcement, and the W3C HAR specification.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790699358) } [4]=> array(11) { ["title"]=> string(79) "The Mandarin Slip: Why ChatGPT, Claude and DeepSeek Sometimes Answer in Chinese" ["link"]=> string(120) "https://nocarolinachronicle.com/mandarinslip-why-chatgpt-claude-deepseek-output-chinese-chain-of-thought-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Tue, 29 Sep 2026 14:17:40 +0000" ["category"]=> string(143) "NewsAnthropicchain of thoughtChatGPTClaudeDeepSeek-R1language confusionlanguage mixingOpenAIOpenAI o1Qwenreasoning modelsreinforcement learning" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53042" ["description"]=> string(553) "Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without rewarding staying in one language, made worse by outcome-based reasoning training. DeepSeek fixed R1's language mixing with an explicit reward; a 14-model study found illegible reasoning in every model but Claude; OpenAI never explained o1's switches; and Anthropic disclosed a 2025 bug that put Thai and Chinese characters in Claude's English answers." ["content"]=> array(1) { ["encoded"]=> string(7316) "

By Bill Moran

If you work long enough with AI, you’ll see output in some strange languages. Even American lab models like ChatGPT and Claude drop into Mandarin, with Chinese characters in the answer and especially in the chain of thought. Novel Cognition went through the research, the postmortems and years of community reports to find out why.

The best-supported answer isn’t that models secretly think in Mandarin. It’s that training rewards correct, useful output and, unless it separately rewards staying in one language, nothing keeps a long answer or a reasoning scratchpad in the language it started in. Mandarin is the switch people notice. It isn’t the only one.

Right answer, wrong language

“Answer this correctly” and “do every step in English” are different objectives, and most training only rewards the first. The Language Confusion Benchmark caught models switching languages across 15 languages, at the level of whole responses, lines and single words; harder prompts and higher temperature made it worse. An ACL 2025 study found ordinary training doesn’t reliably penalize mixed-language text, and that preference tuning which explicitly rejected it fixed much of the problem.

DeepSeek gave the cleanest evidence. DeepSeek-R1-Zero, trained with reinforcement learning on answers alone, began “combining English and Chinese within a single chain-of-thought response.” For R1, DeepSeek added a reward for the share of reasoning written in the target language. Without it, consistency deteriorated during training; with it, consistency held, at a slight cost in reasoning performance. The detail is at Nothing Tells It to Stay in English and The Reward That Fixed DeepSeek-R1.

Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without

The full walkthrough — why training doesn’t penalize language mixing, the reward DeepSeek added to fix R1, the 14-model study that found illegible reasoning everywhere except Claude, OpenAI o1’s unexplained switches, and the Claude bug Anthropic disclosed. Seven minutes.

The scratchpad slips first

A study of 14 reasoning models by Arun Jose (arXiv 2510.27338) found that outcome-trained models such as DeepSeek-R1 and QwQ drift into compressed fragments, unrelated words and non-English characters mid-reasoning, then return to perfectly readable final answers. The abstract names one exception: Claude. Forced to use only the legible portions of their reasoning, the models’ accuracy fell 53%, but the author found no general link between illegibility and correctness and considered a secret language unlikely.

Mandarin dominates the screenshots mostly because it’s visible. OpenAI o1 users also saw its reasoning slip into Persian, Hindi and Thai, and the popular token-efficiency theory isn’t supported by tokenizer research. See Why the Scratchpad Slips First and Why Mandarin, of All Languages.

o1, GPT-5 and Claude do it too

OpenAI never published a cause for o1’s switches into Chinese and Persian; the theory blaming Chinese data-labeling vendors was never demonstrated. Later work found similar degraded reasoning text in o3 and GPT-5, but OpenAI doesn’t expose raw chains of thought. Claude is the one American case with a disclosed cause: Anthropic’s 17 September 2025 postmortem says a TPU server misconfiguration occasionally produced “Thai or Chinese characters in response to English prompts,” affecting Opus 4 and 4.1 from 25 to 28 August and Sonnet 4 until 2 September 2025.

That case is a reminder to check the plumbing first. DeepSeek-Coder-V2’s Chinese output on Ollama traced to a broken chat template; Qwen2.5-VL’s came from padding under batching. The full breakdown is at o1, GPT-5 and Claude Do It Too and Check the Template Before You Blame the Model.

Trainable, not mysterious

Frontier models are multilingual text generators without sealed language modes. Standard training doesn’t reliably penalize accidental language mixing, outcome-based reasoning training makes it worse, and explicit language rewards, in-language examples, preference tuning and careful decoding reduce it. For American models the behavior is documented and, apart from Claude’s disclosed bug, largely undiagnosed. Even the visible chain of thought isn’t a faithful transcript: Anthropic found Claude 3.7 Sonnet mentioned answer-changing hints 25% of the time and DeepSeek-R1 39%.

The full file is at mandarinslip.novcog.us.com. Primary sources: the DeepSeek-R1 paper, Reasoning Models Sometimes Output Illegible Chains of Thought, the Language Confusion Benchmark, and Anthropic’s postmortem.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(553) "Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without rewarding staying in one language, made worse by outcome-based reasoning training. DeepSeek fixed R1's language mixing with an explicit reward; a 14-model study found illegible reasoning in every model but Claude; OpenAI never explained o1's switches; and Anthropic disclosed a 2025 bug that put Thai and Chinese characters in Claude's English answers." ["atom_content"]=> string(7316) "

By Bill Moran

If you work long enough with AI, you’ll see output in some strange languages. Even American lab models like ChatGPT and Claude drop into Mandarin, with Chinese characters in the answer and especially in the chain of thought. Novel Cognition went through the research, the postmortems and years of community reports to find out why.

The best-supported answer isn’t that models secretly think in Mandarin. It’s that training rewards correct, useful output and, unless it separately rewards staying in one language, nothing keeps a long answer or a reasoning scratchpad in the language it started in. Mandarin is the switch people notice. It isn’t the only one.

Right answer, wrong language

“Answer this correctly” and “do every step in English” are different objectives, and most training only rewards the first. The Language Confusion Benchmark caught models switching languages across 15 languages, at the level of whole responses, lines and single words; harder prompts and higher temperature made it worse. An ACL 2025 study found ordinary training doesn’t reliably penalize mixed-language text, and that preference tuning which explicitly rejected it fixed much of the problem.

DeepSeek gave the cleanest evidence. DeepSeek-R1-Zero, trained with reinforcement learning on answers alone, began “combining English and Chinese within a single chain-of-thought response.” For R1, DeepSeek added a reward for the share of reasoning written in the target language. Without it, consistency deteriorated during training; with it, consistency held, at a slight cost in reasoning performance. The detail is at Nothing Tells It to Stay in English and The Reward That Fixed DeepSeek-R1.

Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without

The full walkthrough — why training doesn’t penalize language mixing, the reward DeepSeek added to fix R1, the 14-model study that found illegible reasoning everywhere except Claude, OpenAI o1’s unexplained switches, and the Claude bug Anthropic disclosed. Seven minutes.

The scratchpad slips first

A study of 14 reasoning models by Arun Jose (arXiv 2510.27338) found that outcome-trained models such as DeepSeek-R1 and QwQ drift into compressed fragments, unrelated words and non-English characters mid-reasoning, then return to perfectly readable final answers. The abstract names one exception: Claude. Forced to use only the legible portions of their reasoning, the models’ accuracy fell 53%, but the author found no general link between illegibility and correctness and considered a secret language unlikely.

Mandarin dominates the screenshots mostly because it’s visible. OpenAI o1 users also saw its reasoning slip into Persian, Hindi and Thai, and the popular token-efficiency theory isn’t supported by tokenizer research. See Why the Scratchpad Slips First and Why Mandarin, of All Languages.

o1, GPT-5 and Claude do it too

OpenAI never published a cause for o1’s switches into Chinese and Persian; the theory blaming Chinese data-labeling vendors was never demonstrated. Later work found similar degraded reasoning text in o3 and GPT-5, but OpenAI doesn’t expose raw chains of thought. Claude is the one American case with a disclosed cause: Anthropic’s 17 September 2025 postmortem says a TPU server misconfiguration occasionally produced “Thai or Chinese characters in response to English prompts,” affecting Opus 4 and 4.1 from 25 to 28 August and Sonnet 4 until 2 September 2025.

That case is a reminder to check the plumbing first. DeepSeek-Coder-V2’s Chinese output on Ollama traced to a broken chat template; Qwen2.5-VL’s came from padding under batching. The full breakdown is at o1, GPT-5 and Claude Do It Too and Check the Template Before You Blame the Model.

Trainable, not mysterious

Frontier models are multilingual text generators without sealed language modes. Standard training doesn’t reliably penalize accidental language mixing, outcome-based reasoning training makes it worse, and explicit language rewards, in-language examples, preference tuning and careful decoding reduce it. For American models the behavior is documented and, apart from Claude’s disclosed bug, largely undiagnosed. Even the visible chain of thought isn’t a faithful transcript: Anthropic found Claude 3.7 Sonnet mentioned answer-changing hints 25% of the time and DeepSeek-R1 39%.

The full file is at mandarinslip.novcog.us.com. Primary sources: the DeepSeek-R1 paper, Reasoning Models Sometimes Output Illegible Chains of Thought, the Language Confusion Benchmark, and Anthropic’s postmortem.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790691460) } [5]=> array(11) { ["title"]=> string(98) "The Three Numbers: Qwen 3.8 Fits in 24 GB, Runs 2.24x Faster, and Loses 2.1 Points When Uncensored" ["link"]=> string(105) "https://nocarolinachronicle.com/threenumbers-qwen-3-8-27b-24gb-mtp-2-24x-abliteration-tax-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Mon, 28 Sep 2026 18:25:27 +0000" ["category"]=> string(132) "NewsabliterationApple SiliconKV cachelocal AIMLXMTPLXmulti-token predictionoMLXQwen 3.8Qwen3.8-27Bspeculative decodinguncensored LLM" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53039" ["description"]=> string(452) "In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 layers never cache a token. 2.24x: MTPLX's multi-token prediction on an M5 Max, with the output distribution unchanged. 2.1 points: the MMLU cost of abliteration, published by exactly one of four builds. Every figure is attributed to whoever measured it." ["content"]=> array(1) { ["encoded"]=> string(6624) "

By Bill Moran

Twenty-four gigabytes. 2.24x. 2.1 points. Those three numbers came out of the local AI world inside five weeks, and together they describe something that used to need a data centre. Qwen 3.8 27B, a native vision-language model that holds a quarter of a million tokens, now runs in the memory of a laptop. It got twice as fast, with its output distribution mathematically unchanged. And its ability to refuse was removed with linear algebra.

Two of those numbers are quoted everywhere. The third is what the removal cost, and exactly one builder published it. This is a report on all three, with every figure attributed to whoever measured it.

Why 24 GB runs a 262K window

Qwen 3.8 27B has sixty-four layers. Forty-eight are Gated DeltaNet linear-attention layers holding a fixed recurrent state of about 72 MiB in total; only sixteen are full-attention layers that keep a key-value cache growing with every token. At 64 KiB per token, the full 262,144-token window costs 16 GiB of cache. A conventional all-attention model of the same depth would need 64.

At six-bit weights and a normal working context the model runs in about 24 GB; eight-bit is about 30, and the gap is a flat six gigabytes at every context length, all of it weights. On Apple Silicon, decode speed is set by memory bandwidth, so the larger build is also the slower one. The arithmetic, with the runtime tables, is at Why 24 GB Runs a 262K Window.

In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 la

The full walkthrough — the KV-cache arithmetic behind the 24 GB figure, multi-token prediction and the open regression nobody can find, the one benchmark that measured the abliteration tax, and the install instruction that does not work. Eight minutes.

Twice as fast, same distribution

The model was trained with a multi-token prediction head: it drafts several of its own next tokens and a runtime verifies the whole block in one pass. MTPLX, built for Apple Silicon, commits those tokens through exact rejection sampling with residual correction, so the fast output and the normal output come from the same distribution. It measures 1.6x on a 16 GB M4 Mac mini and 2.24x on an M5 Max. oMLX 0.6.1, released 17 August, measured +34% decode throughput at 16K context with its own Lightning MTP.

The field is far enough into this that an open mlx-vlm issue is hunting a 2.19 ms (+10.6%) regression in MTP verify cycles that nobody has yet localised. Details at Twice as Fast, Same Distribution.

The number almost nobody publishes

Abliteration identifies the direction a model’s internal state moves along when it refuses and orthogonalises that direction out of the residual stream. Every abliterated build of this model reports the same result: zero refusals, or zero out of a hundred. One build, OBLITERATED V3, published on 25 August by the builder working as OBLITERATUS, also measured what the removal cost. MMLU fell from 84.5% to 82.3%. STEM fell from 81.8% to 78.5%. Humanities moved by one point.

That is a real, bounded cost: not catastrophic, not free, and measurable by anyone willing to run the harness. Four builds publish their refusal number to the decimal; one publishes the price. The table is at What Removing Refusals Costs, the mechanism at One Vector, and the four model cards compared at Reading the Model Cards.

What was not tested

Novel Cognition did not re-run MMLU and did not benchmark on its own hardware; every capability and speed figure here is the named builder’s own, reproduced as published. oMLX 0.6.3’s feature list comes from a user post. Whether the fast speculative-decoding paths work with a third-party abliterated drafter has not been shown by anyone. That page is load-bearing and is at What We Did Not Test.

The full file is at threenumbers.novcog.us.com, including the install instruction that does not work. Primary sources: the Qwen3.8-27B model card, MTPLX, the oMLX 0.6.1 release, and OBLITERATED V3.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(452) "In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 layers never cache a token. 2.24x: MTPLX's multi-token prediction on an M5 Max, with the output distribution unchanged. 2.1 points: the MMLU cost of abliteration, published by exactly one of four builds. Every figure is attributed to whoever measured it." ["atom_content"]=> string(6624) "

By Bill Moran

Twenty-four gigabytes. 2.24x. 2.1 points. Those three numbers came out of the local AI world inside five weeks, and together they describe something that used to need a data centre. Qwen 3.8 27B, a native vision-language model that holds a quarter of a million tokens, now runs in the memory of a laptop. It got twice as fast, with its output distribution mathematically unchanged. And its ability to refuse was removed with linear algebra.

Two of those numbers are quoted everywhere. The third is what the removal cost, and exactly one builder published it. This is a report on all three, with every figure attributed to whoever measured it.

Why 24 GB runs a 262K window

Qwen 3.8 27B has sixty-four layers. Forty-eight are Gated DeltaNet linear-attention layers holding a fixed recurrent state of about 72 MiB in total; only sixteen are full-attention layers that keep a key-value cache growing with every token. At 64 KiB per token, the full 262,144-token window costs 16 GiB of cache. A conventional all-attention model of the same depth would need 64.

At six-bit weights and a normal working context the model runs in about 24 GB; eight-bit is about 30, and the gap is a flat six gigabytes at every context length, all of it weights. On Apple Silicon, decode speed is set by memory bandwidth, so the larger build is also the slower one. The arithmetic, with the runtime tables, is at Why 24 GB Runs a 262K Window.

In five weeks, Qwen 3.8 27B moved onto a desk, got twice as fast and had its refusals removed. 24 GB: the native vision-language model runs a 262K-token window in laptop memory because 48 of its 64 la

The full walkthrough — the KV-cache arithmetic behind the 24 GB figure, multi-token prediction and the open regression nobody can find, the one benchmark that measured the abliteration tax, and the install instruction that does not work. Eight minutes.

Twice as fast, same distribution

The model was trained with a multi-token prediction head: it drafts several of its own next tokens and a runtime verifies the whole block in one pass. MTPLX, built for Apple Silicon, commits those tokens through exact rejection sampling with residual correction, so the fast output and the normal output come from the same distribution. It measures 1.6x on a 16 GB M4 Mac mini and 2.24x on an M5 Max. oMLX 0.6.1, released 17 August, measured +34% decode throughput at 16K context with its own Lightning MTP.

The field is far enough into this that an open mlx-vlm issue is hunting a 2.19 ms (+10.6%) regression in MTP verify cycles that nobody has yet localised. Details at Twice as Fast, Same Distribution.

The number almost nobody publishes

Abliteration identifies the direction a model’s internal state moves along when it refuses and orthogonalises that direction out of the residual stream. Every abliterated build of this model reports the same result: zero refusals, or zero out of a hundred. One build, OBLITERATED V3, published on 25 August by the builder working as OBLITERATUS, also measured what the removal cost. MMLU fell from 84.5% to 82.3%. STEM fell from 81.8% to 78.5%. Humanities moved by one point.

That is a real, bounded cost: not catastrophic, not free, and measurable by anyone willing to run the harness. Four builds publish their refusal number to the decimal; one publishes the price. The table is at What Removing Refusals Costs, the mechanism at One Vector, and the four model cards compared at Reading the Model Cards.

What was not tested

Novel Cognition did not re-run MMLU and did not benchmark on its own hardware; every capability and speed figure here is the named builder’s own, reproduced as published. oMLX 0.6.3’s feature list comes from a user post. Whether the fast speculative-decoding paths work with a third-party abliterated drafter has not been shown by anyone. That page is load-bearing and is at What We Did Not Test.

The full file is at threenumbers.novcog.us.com, including the install instruction that does not work. Primary sources: the Qwen3.8-27B model card, MTPLX, the oMLX 0.6.1 release, and OBLITERATED V3.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790619927) } [6]=> array(11) { ["title"]=> string(78) "Cheap Offense: A $3,000 Test Reached Inside OpenAI, and a Model Release Is Why" ["link"]=> string(118) "https://nocarolinachronicle.com/cheapoffense-3000-dollar-authorized-openai-test-model-jump-defense-gap-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Sat, 26 Sep 2026 20:46:21 +0000" ["category"]=> string(130) "NewsAI capabilitiesAI policyAI securityAnthropicbug bountyCISAClaudeClaude Opus 5CybersecurityHacktronOpenAIresponsible disclosure" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53036" ["description"]=> string(563) "Three researchers at Hacktron AI ran an authorized bug-bounty test and reached OpenAI's internal systems; OpenAI paid $6,500 and fixed it in about 14 hours. The story is not the hack. The whole two-month project cost under $3,000 in AI tokens, its hardest step succeeded only once Claude Opus 5 shipped mid-project, and the model refused the task until it was reworded. Meanwhile the US civilian cyber-defense agency, CISA, has lost about a third of its staff. Offense got cheap while public defense got cut. Reported at the journalism layer, no method described." ["content"]=> array(1) { ["encoded"]=> string(6959) "

By Bill Moran

Three security researchers reached inside OpenAI this summer. Not as criminals, and not by accident. Harsh Jaiswal, Mohan Pedhapati and Rahul Maini of the security firm Hacktron AI ran an authorized test under OpenAI’s own bug-bounty program, proved they could reach an internal code repository with a single harmless action, and stopped. OpenAI confirmed a fix within about fourteen hours and paid a $6,500 bounty. This is how responsible disclosure is supposed to work.

The unsettling part is not the break. It is the receipt. The entire two-month project cost the team under $3,000 in AI tokens, and its hardest step only succeeded because a newer AI model shipped in the middle of the work. This is a report on what that means, at the level of cost, capability and policy. No method appears here by choice.

A version number did what effort could not

Strip out the target and the reward and one detail is the story. The hardest part of the work did not yield to effort. It yielded to a release. For weeks, Claude Opus 4.8 could not finish the key step across repeated sessions. Then Opus 5 shipped, mid-project, and in the researchers’ own words a task the older model had failed across several sessions was solved by the newer one within hours of its release.

“AI can hack” has been a headline for two years. This is sharper. The capability frontier for real offensive work is moving in visible steps, and each step is a scheduled product launch. A defender who was safe against the tooling in June was not safe in late July, and nothing on their end changed. Their exposure moved on a calendar they do not control. The full walk-through is at The Model Jump.

Three researchers at Hacktron AI ran an authorized bug-bounty test and reached OpenAI’s internal systems; OpenAI paid $6,500 and fixed it in about 14 hours. The story is not the hack. The whole two-mo

The full analysis — the authorized test, the model-generation jump that made it cheap, the guardrail that held only until it was reworded, and the defensive agency being cut at the same time. Seven minutes.

When offense costs less than a laptop

The bounty was $6,500. The project cost under $3,000 in tokens. Security has quietly relied on cost as a control for decades: the assumption that work this deep needed a funded team and a long runway, so only serious, rare actors could afford it. That assumption broke this summer, and it broke quietly.

When the price of an attempt falls by an order of magnitude, more people can try and each skilled person can run more attempts at once. The threat model built for a small number of expensive actors does not survive a large number of cheap ones. The point is not panic. It is to stop pricing risk on last year’s cost of an attack. The full reasoning is at Under Three Thousand, and the guardrail question — a refusal that held until it was reworded — is at The Refusal.

Offense got cheap. Public defense got cut.

This becomes a public-policy story when you set it beside the other trend line. As AI-assisted offense grew cheaper in 2026, the country’s civilian cyber-defense agency was thinned. CISA has lost roughly a third of its workforce since January 2025 and has had no permanent director in that time. The budget proposed for the next year would cut it further — on the order of hundreds of millions of dollars and hundreds of positions, including a roughly 60 percent cut to the work of running government penetration tests.

There is, in fairness, a Trump administration executive order aimed at roughly this, a framework for early government access to test frontier models. But an order on paper does not close a gap; funded people do, and the proposal is to have fewer of them. The figures, with both endpoints and both framings, are at The Defense Gap.

Price your risk on this year’s cost

The event itself was good-faith work by skilled people, disclosed and fixed, and the researchers handled it well. The lesson is not about them. It is that the floor for this kind of work dropped hard this year, that a single model release can move it again without warning, and that the safety net most organizations quietly count on is being cut rather than grown.

The full analysis, with both ends of every date and every figure sourced, is at cheapoffense.novcog.us.com: What Happened, The Model Jump, The Defense Gap, and Sources & Method. Primary source: the researchers’ own writeup at hacktron.ai; corroborated by The Register, Malwarebytes and The Next Web.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(563) "Three researchers at Hacktron AI ran an authorized bug-bounty test and reached OpenAI's internal systems; OpenAI paid $6,500 and fixed it in about 14 hours. The story is not the hack. The whole two-month project cost under $3,000 in AI tokens, its hardest step succeeded only once Claude Opus 5 shipped mid-project, and the model refused the task until it was reworded. Meanwhile the US civilian cyber-defense agency, CISA, has lost about a third of its staff. Offense got cheap while public defense got cut. Reported at the journalism layer, no method described." ["atom_content"]=> string(6959) "

By Bill Moran

Three security researchers reached inside OpenAI this summer. Not as criminals, and not by accident. Harsh Jaiswal, Mohan Pedhapati and Rahul Maini of the security firm Hacktron AI ran an authorized test under OpenAI’s own bug-bounty program, proved they could reach an internal code repository with a single harmless action, and stopped. OpenAI confirmed a fix within about fourteen hours and paid a $6,500 bounty. This is how responsible disclosure is supposed to work.

The unsettling part is not the break. It is the receipt. The entire two-month project cost the team under $3,000 in AI tokens, and its hardest step only succeeded because a newer AI model shipped in the middle of the work. This is a report on what that means, at the level of cost, capability and policy. No method appears here by choice.

A version number did what effort could not

Strip out the target and the reward and one detail is the story. The hardest part of the work did not yield to effort. It yielded to a release. For weeks, Claude Opus 4.8 could not finish the key step across repeated sessions. Then Opus 5 shipped, mid-project, and in the researchers’ own words a task the older model had failed across several sessions was solved by the newer one within hours of its release.

“AI can hack” has been a headline for two years. This is sharper. The capability frontier for real offensive work is moving in visible steps, and each step is a scheduled product launch. A defender who was safe against the tooling in June was not safe in late July, and nothing on their end changed. Their exposure moved on a calendar they do not control. The full walk-through is at The Model Jump.

Three researchers at Hacktron AI ran an authorized bug-bounty test and reached OpenAI’s internal systems; OpenAI paid $6,500 and fixed it in about 14 hours. The story is not the hack. The whole two-mo

The full analysis — the authorized test, the model-generation jump that made it cheap, the guardrail that held only until it was reworded, and the defensive agency being cut at the same time. Seven minutes.

When offense costs less than a laptop

The bounty was $6,500. The project cost under $3,000 in tokens. Security has quietly relied on cost as a control for decades: the assumption that work this deep needed a funded team and a long runway, so only serious, rare actors could afford it. That assumption broke this summer, and it broke quietly.

When the price of an attempt falls by an order of magnitude, more people can try and each skilled person can run more attempts at once. The threat model built for a small number of expensive actors does not survive a large number of cheap ones. The point is not panic. It is to stop pricing risk on last year’s cost of an attack. The full reasoning is at Under Three Thousand, and the guardrail question — a refusal that held until it was reworded — is at The Refusal.

Offense got cheap. Public defense got cut.

This becomes a public-policy story when you set it beside the other trend line. As AI-assisted offense grew cheaper in 2026, the country’s civilian cyber-defense agency was thinned. CISA has lost roughly a third of its workforce since January 2025 and has had no permanent director in that time. The budget proposed for the next year would cut it further — on the order of hundreds of millions of dollars and hundreds of positions, including a roughly 60 percent cut to the work of running government penetration tests.

There is, in fairness, a Trump administration executive order aimed at roughly this, a framework for early government access to test frontier models. But an order on paper does not close a gap; funded people do, and the proposal is to have fewer of them. The figures, with both endpoints and both framings, are at The Defense Gap.

Price your risk on this year’s cost

The event itself was good-faith work by skilled people, disclosed and fixed, and the researchers handled it well. The lesson is not about them. It is that the floor for this kind of work dropped hard this year, that a single model release can move it again without warning, and that the safety net most organizations quietly count on is being cut rather than grown.

The full analysis, with both ends of every date and every figure sourced, is at cheapoffense.novcog.us.com: What Happened, The Model Jump, The Defense Gap, and Sources & Method. Primary source: the researchers’ own writeup at hacktron.ai; corroborated by The Register, Malwarebytes and The Next Web.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790455581) } [7]=> array(11) { ["title"]=> string(117) "He Sells Courses to $79,995 on a Guarantee of Business Success. Neither Website Publishes What a Typical Buyer Earns." ["link"]=> string(106) "https://nocarolinachronicle.com/marshall-sylver-prosperity-alliance-seminar-claims-checked-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Tue, 22 Sep 2026 18:12:20 +0000" ["category"]=> string(166) "Newsbusiness coachingbusiness opportunityconsumer protectionearnings claimsFTCinvestigationLas VegasMarshall SylverProsperity Alliancerefund policyseminartestimonials" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53033" ["description"]=> string(1579) "
Comparison table showing prosperityalliance.com and sylver.com lack typical earnings disclosures for $79,995 programs.
Marshall Sylver's Prosperity Alliance sells a ladder of courses from $1,995 to $79,995 on the published promise that success in almost any business is guaranteed. Neither of his websites publishes an earnings disclosure, and one publishes no terms or refund policy at all. Identical testimonial text appears under different names on his own pages - in one pair, carrying the same misspelling. He was indicted in 2003 and never convicted; no federal regulator has ever acted against him." ["content"]=> array(1) { ["encoded"]=> string(16217) "
Comparison table showing prosperityalliance.com and sylver.com lack typical earnings disclosures for $79,995 programs.

By Bill Moran

Prosperity Alliance employs a tiered pricing model for its wealth-building courses, with entry-level instruction starting at $1,995 and premium packages reaching $79,995. These high-ticket offerings are marketed with guarantees of business success and the promise of a first million dollars. An investigation by the Keystone Gazette examined the actual return on investment for these expenditures, contrasting the company’s published claims against the documented record.

The civil record

Sylver has appeared in court often, on both sides of the caption. None of the following is a finding of dishonesty, and none is presented as one.

In 2001 the Venetian sued him over a casino credit marker, claiming roughly $201,000 outstanding on a $205,000 marker. Sylver disputed it at the time: “There’s some confusion over the financing of the marker. The Venetian has been receiving payments on a regular basis. I’ve been going through a divorce since January 2000.” The case was reported dismissed in 2004.

That same year he sued a former customer, Sean Roach, alleging a “malicious campaign of lies, stalking and personal attacks.” The Las Vegas Sun reported that Roach was also one of the customers who had complained about him to the attorney general.

In Sylver v. Regents Bank, decided by the Nevada Supreme Court in 2013, he sought to vacate an arbitration award over two 2008 bridge loans. The court affirmed against him, holding that “Sylver has not met his burden for vacating the arbitration award.”

In January 2013 he announced a plan to take over PH Live, a 7,000-seat Las Vegas theatre, in a venture he valued at $400 million. On 24 June 2013 — four months later — he filed for personal Chapter 11 bankruptcy. A discharge was entered in March 2023.

He has also sued and lost. A federal court granted summary judgment against him in a fraud action he brought in 2009. In another, he sued a charter operator over being supplied a Challenger 601 rather than a Gulfstream 200, pleading among other things a violation of the Nevada Deceptive Trade Practices Act — the statute the attorney general had cited against him nine years earlier.

Asked about litigation in 2013, Sylver told the Las Vegas Sun: “If you are in my type of business long enough, you will wind up in litigation.”

Prosperity Alliance’s website publishes no terms of service, no privacy policy and no refund policy.

“There’s no risks, no small print”

On 30 October 2001 the Nevada attorney general’s office announced a criminal investigation and served a search warrant on Sylver’s home office. Its own press release, issued under Attorney General Frankie Sue Del Papa, said the office “suspected that Sylver has committed the criminal offenses of Theft by Obtaining Money Under False Pretenses, a felony; Racketeering, a felony; and misdemeanor violations of the Deceptive Trade Practices Act.”

That release ended with a sentence worth repeating in full: “As in all criminal matters, the allegations are merely accusations and individuals are presumed innocent unless and until proven guilty in court.”

A Clark County grand jury returned an indictment in April 2003. The attorney general’s announcement set out the allegations: that Sylver “refused to honor his promised money-back guarantee, that he did not provide the promised mentoring and that he did not pre-interview anyone” for the Millionaire Mentorship Program.

The programme cost between $4,500 and $6,500 for a three-and-a-half day seminar and ten weeks of mentoring. As prosecutors rested in December 2003, they played Sylver’s radio advertisement to the jury twice. It said: “There’s no risks, no small print. Do what I tell you to do and you’ll be on your way to becoming a multimillionaire.”

The guarantee that advertisement promoted — money back if a client did not double their investment in ninety days — carried three written conditions. Attend every class. Speak with a mentor every weekday. Complete every daily assignment. The defence at trial was that the complainants had not met them.

The company’s own director of mentor services testified to the standard applied: “If you miss a few letters of the alphabet, you didn’t correctly say it.” He further testified that after refund requests began arriving in late 2000, the company introduced a $1,000 “personal responsibility discount” offered to new customers who agreed to waive the money-back guarantee entirely.

Nine clients complained, out of a programme with roughly 1,200. Five of them recovered their money separately, by winning judgments in small claims court. After about four days of deliberation the jury deadlocked and District Judge Valorie Vega declared a mistrial. Prosecutors said at the time that a retrial was likely. It never happened.

The documents that are not there

A seller making a specific claim about money would ordinarily be able to point to the basis for it. Prosperity Alliance’s website publishes no terms of service, no privacy policy and no refund policy — a finding confirmed against a crawl of 150 URLs rather than by guessing at addresses. The page carrying the entire price ladder is served with an instruction asking search engines not to index it.

A companion site, sylver.com, does publish terms. It publishes two refund positions, and they contradict each other. The Terms of Use state that the purchaser’s “exclusive remedy” is a refund of the price paid, “typically limit[ed] to 30 days.” The private coaching page on the same website states: “All Sales are final… No refunds will be given.” The order page for a live seminar advertised for October 2026 carries no refund terms at all.

Neither site publishes an earnings disclosure — the document that would state what proportion of purchasers achieved the advertised outcome. The company may hold such material. A prospective buyer cannot read it before paying.

That absence sits beside specific figures. The testimonial page carries claims of “$2,500,000 in first year”, “I made $10,000 the very first week” and “In four weeks I earned over $40,000.”

The same testimonial, under different names

Identical testimonial text appears on Sylver’s own websites attributed to different named individuals. There are five documented instances; three of them sit on a single page.

On the Prosperity Alliance home page, a sentence praising the seminar and “the fire eating” appears twice, attributed to two different people. On the testimonial page, a thank-you note about “unlimited re-attends” appears twice, again under two names.

The spelling is what makes the duplication difficult to explain. In one pair, the misspelling “Tuning Point” — for Turning Point — appears in both copies. In the other, one copy reads “friends and acquientances” and the second reads “acquaintances”, correctly: the same sentence, two names, the typographical error corrected in one of them.

This newspaper does not assert that any individual review is fabricated; it has no way of knowing who wrote them. The finding is duplication, and it is verifiable by anyone willing to compare two pages.

The testimonial page carries exactly nine customer reviews. All nine are rated five stars. All nine are dated between 16 January and 24 April 2017 — a window of roughly fourteen weeks — with none before and none since in the nine years the page has been online. The site’s terms state that “all testimonials appear after they have been reviewed by management.”

Evidence pointing the other way is also on the record. Sylver’s Trustpilot profile carries three reviews and a 2.8 rating. Trustpilot’s own transparency data shows the profile has never been claimed, that no review invitations were ever sent, and that no review has ever been flagged. There is no evidence of manipulation on that platform, and none is alleged. The duplication occurs entirely within marketing material his companies publish themselves.

The business address is a mailbox at a shipping store

The Better Business Bureau lists an address for Prosperity Alliance: 1027 S Rainbow Blvd # 281, Las Vegas, NV 89145. It lists the same address, with the same box number, for a second Sylver company, Mind Power, Inc.

That address is The UPS Store #1267, in Rainbow Plaza next to an Albertsons supermarket.

The suite number is the tell. The UPS Store’s own mailbox-services page sets out the format a customer receives when they rent a private mailbox there: “Joe Smith PMB XXX or # XXX, 1027 S Rainbow Blvd, Las Vegas, NV 89145.” The “# 281” on both business records is that format exactly.

Renting a private mailbox is entirely lawful, and plenty of legitimate small businesses do it. Two things still make it worth reporting here. The first is that two separate corporate entities share a single box. The second is the distance between the address and the offer: the top tier of the course ladder, at $79,995, is sold as time at the seller’s “Personal Oasis, The Prosperity Palace,” while the business behind it receives its post in a strip mall.

A buyer weighing a five-figure purchase is entitled to know where the company they are paying can actually be found.

The dismissal that cannot be sourced

Search for this case today and most summaries state that it was dismissed in 2005 and the charges withdrawn. The Keystone Gazette does not print that, and the reason is itself a finding.

The Nevada attorney general’s bound press archives from 1998 to 2011 were searched in full text. They contain exactly two items mentioning Sylver in fourteen years: the 2001 raid and the 2003 indictment. No release announces a dismissal, a retrial, a withdrawal, a conviction or an acquittal. Neither the Las Vegas Weekly in 2009 nor the Las Vegas Sun in 2013 mentions a dismissal; both stop at the mistrial.

The sentence survives in one place: a Wikipedia revision that was deleted, now preserved only on a mirror site. Its sole citation is a print gossip column from June 2007. The archived talk page shows the material was added during a dispute in which other editors believed the contributing account was connected to Sylver.

That deleted text now circulates through AI-generated encyclopaedias and into search summaries, where it reads as settled fact. What is established is narrower: he was indicted, tried, the jury deadlocked, prosecutors said they intended to retry him, he was never retried, and he was never convicted.

Why the refund is not the protection

None of the federal enforcement below names Sylver or any of his companies, and none is alleged to. It establishes what this category looks like when a regulator does reach a finding.

In December 2020 the Federal Trade Commission announced Operation Income Illusion, a sweep with nineteen federal, state and local partners comprising more than fifty law enforcement actions. The Commission’s own description of what it targeted includes, in its words, “bogus coaching courses.” Consumers had reported losing more than $610 million to such schemes since 2016.

In one case folded into that sweep, monetary judgments came to more than $32 million. Those judgments were partially suspended on the defendants’ inability to pay. The assets actually surrendered came to just over $1.25 million — roughly four cents on the dollar.

In September 2025 the Commission announced it was distributing $666,631 to 4,208 consumers harmed by a business-opportunity scheme, about $158 each, against a category that sells in the thousands.

In January 2025 the FTC proposed extending its substantiation requirements to “money-making opportunities” generally, defined to include business coaching. Covered sellers would have to hold written substantiation for earnings claims and provide it to consumers on request. Sam Levine, then director of the Bureau of Consumer Protection, put the harm plainly: “Phony claims about likely earnings lure people looking for honest income into spending thousands, even tens of thousands, of dollars.”

That is the document this newspaper looked for and did not find. A buyer does not need to wait for a rule to ask for it.

What to ask before you pay

One question does most of the work, and it requires no expertise: what percentage of everyone who paid achieved the result you are advertising? Not the best customer. All of them. Over what period, and measured how?

Ask in writing and keep the answer. The reply matters more than the answer. A seller with evidence produces a figure, a period, a denominator, a document. If what comes back is more testimonials, an invitation to a call, language about mindset or resistance, or a suggestion that the real results are at the next tier — that is the answer.

Marshall Sylver was contacted before publication and did not make substantive comment. The full investigation, every source, and an explicit list of what could not be established — including a claim about his record that this newspaper declined to print — is at nosmallprint.keystonegazette.com.


Part of the Investigations Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(1579) "
Comparison table showing prosperityalliance.com and sylver.com lack typical earnings disclosures for $79,995 programs.
Marshall Sylver's Prosperity Alliance sells a ladder of courses from $1,995 to $79,995 on the published promise that success in almost any business is guaranteed. Neither of his websites publishes an earnings disclosure, and one publishes no terms or refund policy at all. Identical testimonial text appears under different names on his own pages - in one pair, carrying the same misspelling. He was indicted in 2003 and never convicted; no federal regulator has ever acted against him." ["atom_content"]=> string(16217) "
Comparison table showing prosperityalliance.com and sylver.com lack typical earnings disclosures for $79,995 programs.

By Bill Moran

Prosperity Alliance employs a tiered pricing model for its wealth-building courses, with entry-level instruction starting at $1,995 and premium packages reaching $79,995. These high-ticket offerings are marketed with guarantees of business success and the promise of a first million dollars. An investigation by the Keystone Gazette examined the actual return on investment for these expenditures, contrasting the company’s published claims against the documented record.

The civil record

Sylver has appeared in court often, on both sides of the caption. None of the following is a finding of dishonesty, and none is presented as one.

In 2001 the Venetian sued him over a casino credit marker, claiming roughly $201,000 outstanding on a $205,000 marker. Sylver disputed it at the time: “There’s some confusion over the financing of the marker. The Venetian has been receiving payments on a regular basis. I’ve been going through a divorce since January 2000.” The case was reported dismissed in 2004.

That same year he sued a former customer, Sean Roach, alleging a “malicious campaign of lies, stalking and personal attacks.” The Las Vegas Sun reported that Roach was also one of the customers who had complained about him to the attorney general.

In Sylver v. Regents Bank, decided by the Nevada Supreme Court in 2013, he sought to vacate an arbitration award over two 2008 bridge loans. The court affirmed against him, holding that “Sylver has not met his burden for vacating the arbitration award.”

In January 2013 he announced a plan to take over PH Live, a 7,000-seat Las Vegas theatre, in a venture he valued at $400 million. On 24 June 2013 — four months later — he filed for personal Chapter 11 bankruptcy. A discharge was entered in March 2023.

He has also sued and lost. A federal court granted summary judgment against him in a fraud action he brought in 2009. In another, he sued a charter operator over being supplied a Challenger 601 rather than a Gulfstream 200, pleading among other things a violation of the Nevada Deceptive Trade Practices Act — the statute the attorney general had cited against him nine years earlier.

Asked about litigation in 2013, Sylver told the Las Vegas Sun: “If you are in my type of business long enough, you will wind up in litigation.”

Prosperity Alliance’s website publishes no terms of service, no privacy policy and no refund policy.

“There’s no risks, no small print”

On 30 October 2001 the Nevada attorney general’s office announced a criminal investigation and served a search warrant on Sylver’s home office. Its own press release, issued under Attorney General Frankie Sue Del Papa, said the office “suspected that Sylver has committed the criminal offenses of Theft by Obtaining Money Under False Pretenses, a felony; Racketeering, a felony; and misdemeanor violations of the Deceptive Trade Practices Act.”

That release ended with a sentence worth repeating in full: “As in all criminal matters, the allegations are merely accusations and individuals are presumed innocent unless and until proven guilty in court.”

A Clark County grand jury returned an indictment in April 2003. The attorney general’s announcement set out the allegations: that Sylver “refused to honor his promised money-back guarantee, that he did not provide the promised mentoring and that he did not pre-interview anyone” for the Millionaire Mentorship Program.

The programme cost between $4,500 and $6,500 for a three-and-a-half day seminar and ten weeks of mentoring. As prosecutors rested in December 2003, they played Sylver’s radio advertisement to the jury twice. It said: “There’s no risks, no small print. Do what I tell you to do and you’ll be on your way to becoming a multimillionaire.”

The guarantee that advertisement promoted — money back if a client did not double their investment in ninety days — carried three written conditions. Attend every class. Speak with a mentor every weekday. Complete every daily assignment. The defence at trial was that the complainants had not met them.

The company’s own director of mentor services testified to the standard applied: “If you miss a few letters of the alphabet, you didn’t correctly say it.” He further testified that after refund requests began arriving in late 2000, the company introduced a $1,000 “personal responsibility discount” offered to new customers who agreed to waive the money-back guarantee entirely.

Nine clients complained, out of a programme with roughly 1,200. Five of them recovered their money separately, by winning judgments in small claims court. After about four days of deliberation the jury deadlocked and District Judge Valorie Vega declared a mistrial. Prosecutors said at the time that a retrial was likely. It never happened.

The documents that are not there

A seller making a specific claim about money would ordinarily be able to point to the basis for it. Prosperity Alliance’s website publishes no terms of service, no privacy policy and no refund policy — a finding confirmed against a crawl of 150 URLs rather than by guessing at addresses. The page carrying the entire price ladder is served with an instruction asking search engines not to index it.

A companion site, sylver.com, does publish terms. It publishes two refund positions, and they contradict each other. The Terms of Use state that the purchaser’s “exclusive remedy” is a refund of the price paid, “typically limit[ed] to 30 days.” The private coaching page on the same website states: “All Sales are final… No refunds will be given.” The order page for a live seminar advertised for October 2026 carries no refund terms at all.

Neither site publishes an earnings disclosure — the document that would state what proportion of purchasers achieved the advertised outcome. The company may hold such material. A prospective buyer cannot read it before paying.

That absence sits beside specific figures. The testimonial page carries claims of “$2,500,000 in first year”, “I made $10,000 the very first week” and “In four weeks I earned over $40,000.”

The same testimonial, under different names

Identical testimonial text appears on Sylver’s own websites attributed to different named individuals. There are five documented instances; three of them sit on a single page.

On the Prosperity Alliance home page, a sentence praising the seminar and “the fire eating” appears twice, attributed to two different people. On the testimonial page, a thank-you note about “unlimited re-attends” appears twice, again under two names.

The spelling is what makes the duplication difficult to explain. In one pair, the misspelling “Tuning Point” — for Turning Point — appears in both copies. In the other, one copy reads “friends and acquientances” and the second reads “acquaintances”, correctly: the same sentence, two names, the typographical error corrected in one of them.

This newspaper does not assert that any individual review is fabricated; it has no way of knowing who wrote them. The finding is duplication, and it is verifiable by anyone willing to compare two pages.

The testimonial page carries exactly nine customer reviews. All nine are rated five stars. All nine are dated between 16 January and 24 April 2017 — a window of roughly fourteen weeks — with none before and none since in the nine years the page has been online. The site’s terms state that “all testimonials appear after they have been reviewed by management.”

Evidence pointing the other way is also on the record. Sylver’s Trustpilot profile carries three reviews and a 2.8 rating. Trustpilot’s own transparency data shows the profile has never been claimed, that no review invitations were ever sent, and that no review has ever been flagged. There is no evidence of manipulation on that platform, and none is alleged. The duplication occurs entirely within marketing material his companies publish themselves.

The business address is a mailbox at a shipping store

The Better Business Bureau lists an address for Prosperity Alliance: 1027 S Rainbow Blvd # 281, Las Vegas, NV 89145. It lists the same address, with the same box number, for a second Sylver company, Mind Power, Inc.

That address is The UPS Store #1267, in Rainbow Plaza next to an Albertsons supermarket.

The suite number is the tell. The UPS Store’s own mailbox-services page sets out the format a customer receives when they rent a private mailbox there: “Joe Smith PMB XXX or # XXX, 1027 S Rainbow Blvd, Las Vegas, NV 89145.” The “# 281” on both business records is that format exactly.

Renting a private mailbox is entirely lawful, and plenty of legitimate small businesses do it. Two things still make it worth reporting here. The first is that two separate corporate entities share a single box. The second is the distance between the address and the offer: the top tier of the course ladder, at $79,995, is sold as time at the seller’s “Personal Oasis, The Prosperity Palace,” while the business behind it receives its post in a strip mall.

A buyer weighing a five-figure purchase is entitled to know where the company they are paying can actually be found.

The dismissal that cannot be sourced

Search for this case today and most summaries state that it was dismissed in 2005 and the charges withdrawn. The Keystone Gazette does not print that, and the reason is itself a finding.

The Nevada attorney general’s bound press archives from 1998 to 2011 were searched in full text. They contain exactly two items mentioning Sylver in fourteen years: the 2001 raid and the 2003 indictment. No release announces a dismissal, a retrial, a withdrawal, a conviction or an acquittal. Neither the Las Vegas Weekly in 2009 nor the Las Vegas Sun in 2013 mentions a dismissal; both stop at the mistrial.

The sentence survives in one place: a Wikipedia revision that was deleted, now preserved only on a mirror site. Its sole citation is a print gossip column from June 2007. The archived talk page shows the material was added during a dispute in which other editors believed the contributing account was connected to Sylver.

That deleted text now circulates through AI-generated encyclopaedias and into search summaries, where it reads as settled fact. What is established is narrower: he was indicted, tried, the jury deadlocked, prosecutors said they intended to retry him, he was never retried, and he was never convicted.

Why the refund is not the protection

None of the federal enforcement below names Sylver or any of his companies, and none is alleged to. It establishes what this category looks like when a regulator does reach a finding.

In December 2020 the Federal Trade Commission announced Operation Income Illusion, a sweep with nineteen federal, state and local partners comprising more than fifty law enforcement actions. The Commission’s own description of what it targeted includes, in its words, “bogus coaching courses.” Consumers had reported losing more than $610 million to such schemes since 2016.

In one case folded into that sweep, monetary judgments came to more than $32 million. Those judgments were partially suspended on the defendants’ inability to pay. The assets actually surrendered came to just over $1.25 million — roughly four cents on the dollar.

In September 2025 the Commission announced it was distributing $666,631 to 4,208 consumers harmed by a business-opportunity scheme, about $158 each, against a category that sells in the thousands.

In January 2025 the FTC proposed extending its substantiation requirements to “money-making opportunities” generally, defined to include business coaching. Covered sellers would have to hold written substantiation for earnings claims and provide it to consumers on request. Sam Levine, then director of the Bureau of Consumer Protection, put the harm plainly: “Phony claims about likely earnings lure people looking for honest income into spending thousands, even tens of thousands, of dollars.”

That is the document this newspaper looked for and did not find. A buyer does not need to wait for a rule to ask for it.

What to ask before you pay

One question does most of the work, and it requires no expertise: what percentage of everyone who paid achieved the result you are advertising? Not the best customer. All of them. Over what period, and measured how?

Ask in writing and keep the answer. The reply matters more than the answer. A seller with evidence produces a figure, a period, a denominator, a document. If what comes back is more testimonials, an invitation to a call, language about mindset or resistance, or a suggestion that the real results are at the next tier — that is the answer.

Marshall Sylver was contacted before publication and did not make substantive comment. The full investigation, every source, and an explicit list of what could not be established — including a claim about his record that this newspaper declined to print — is at nosmallprint.keystonegazette.com.


Part of the Investigations Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790100740) } [8]=> array(11) { ["title"]=> string(95) "TypeSafe Claims Its Model Is 193.6x Faster. Its Own Employee Measured 15.9% on a Real Pipeline." ["link"]=> string(93) "https://nocarolinachronicle.com/jev-typesafe-benchmark-checked-explainer-wave-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Mon, 21 Sep 2026 20:05:36 +0000" ["category"]=> string(139) "NewsAI benchmarksAI engineeringAI hypecalibrationclassifierJevLLMmachine learningmodel evaluationstructured outputTypeSafevendor benchmarks" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53030" ["description"]=> string(1523) "
Screenshot from TypeSafe evals explaining that reference labels are generated by GPT-6 Astra and Claude Fable 5.1.
A 33-minute explainer calls TypeSafe's Jev classifier one of the biggest events in the history of computer science. Its benchmark defines the correct answer as the average of two frontier LLMs, so it measures agreement rather than accuracy — TypeSafe says so itself. Every published speed measurement lands between 1.16x and 6x against a 193.6x headline, the lowest coming from TypeSafe's own employee. And a limitations section promised three times is absent from the video's own chapter markers." ["content"]=> array(1) { ["encoded"]=> string(10584) "
Screenshot from TypeSafe evals explaining that reference labels are generated by GPT-6 Astra and Claude Fable 5.1.

By Bill Moran

TypeSafe is positioning its new Jev classifier as a disruptive force in computer science, promising a combination of high speed, low cost, and rigorous type guarantees. However, an analysis of the vendor’s claims and a recent 33-minute explainer suggests the actual economic impact may be overstated. While the pricing model is legitimate, the scale of the promised efficiencies and several key promotional claims are not supported by the source data.

What independent testing found

An external benchmark published 17 September ran 2,000 synthetic phishing emails, reproducibly, with no vendor relationship. Asked the straightforward question once — is this phishing — Jev scored 62.6%. Claude Haiku 4.5, a small and inexpensive model that has been shipping for a year, scored 81.3%.

A 95% figure from the same benchmark also circulates. Reaching it required splitting the single question into five narrower ones, hand-labelling 1,000 examples and fitting a logistic regression over the results. As the researchers put it, that 95% is not Jev; it is Jev plus your labelled data plus a regression you maintain — approximately the work a general-purpose classifier was sold to remove.

The most consequential result for anyone planning to route on Jev’s confidence scores came from a pre-registered evaluation on 20 September. Thirty messages falling outside the scope of every available answer were put in front of the model. It flagged none of them, at 0.99 confidence. On a separate unsolvable task it was correct 44.7% of the time while reporting an average probability of 0.74. Out of distribution, its expected calibration error measured 0.107, some 4.4 times the noise floor.

The type-safety guarantee is real but narrow: output cannot fall outside the schema. A forced-choice model cannot answer “none of these” unless that option was supplied, so an out-of-scope input does not produce an error. It produces a confident wrong answer. The model does not invent text. It invents certainty.

The zero is a definition, not a measurement.

Every published measurement lands between 1.16x and 6x

TypeSafe’s home page claims 193.6x faster and 444.6x cheaper. Set against that, the measurements anyone has actually published:

Both the large and the small numbers can be honest at once. The 193.6x figure compares a single decision call against a single frontier-model call on workflows the vendor wrote. The 15.9% figure measures what happens to an entire pipeline when one step inside it is swapped. Replace one component with something four hundred times cheaper and the system does not become four hundred times cheaper; the remaining calls still dominate.

Which is the point. The figure that matters to anyone deciding whether to adopt is the pipeline figure, and the pipeline figure is the one nobody quotes. The video’s presenter cites the most favourable real measurement available — six times — and then endorses a claim of roughly a hundred.

The full walkthrough, nine minutes:

Watch on YouTube, or read the full analysis at fallsover.novcog.us.com.

The benchmark defines the correct answer as what two other models say

Every speed and cost figure in the launch comes from an evaluation TypeSafe built. The company describes the scoring itself, on its own blog: “We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models.”

The consequence is structural. A perfect score on that benchmark would mean agreeing with those two frontier models on every question, including the ones they get wrong. Jev cannot outperform them, because they are the grading key. The number does not measure whether the model is correct about the world; it measures how closely and how cheaply it imitates the models it is being sold as a replacement for.

TypeSafe also records that the workflows “were made by individuals on our model capabilities team, so some bias could exist,” and that competing models were run through TypeSafe’s own wrapper. None of this is hidden. It is simply not carried along when the headline figure travels.

A related figure has the same shape. TypeSafe’s launch chart shows a 0% hallucination rate. The footnote beneath it reads: “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.” The zero is a definition, not a measurement.

A limitations section, promised three times

The video undertakes to explain where Jev falls over on three separate occasions: in the opening line of its description, again in the bulleted contents of that description, and twice in the narration, at 00:40 and 10:59.

What arrives, at 23:00, runs about twenty-five seconds: “Now, is Jev perfect? Of course not. It’s an excellent general purpose classifier. Like anything in computer science, there’s no perfect tool for everything. It still makes some mistakes.” No failure mode is named. No figure is given. Nothing in it is falsifiable.

This does not rest on a characterisation of the video’s structure, because its author published that structure. The chapter markers in the description read: 23:06 Testing Jev and getting started; 26:29 Cost and speed; 27:20 Jevons paradox and the future of software. There is no limitations chapter, because there is no limitations section. The twenty-five seconds above sit inside “Testing Jev and getting started.”

Material for such a section existed by 19 September. A reproducible outside benchmark had been published two days earlier. Calibration problems had been measured. None of it appears.

What survives

Three claims hold up without qualification. Because Jev answers in a single parallel pass rather than generating tokens one at a time, end-to-end latency of 70 to 500 milliseconds follows from the architecture. Because output is constrained to a caller-defined schema, malformed output cannot occur. And the pricing — $0.042 per million input tokens, with output unmetered — recomputes exactly; Novel Cognition checked the arithmetic on the assumption it was wrong, and it is not.

A fast, inexpensive, schema-locked classifier is a useful component, and the economics are worth testing. Three things are worth doing before relying on one: run it against your own labelled data rather than the vendor’s workflows; feed it inputs matching none of your options and observe what comes back; and measure calibration on your own distribution before routing anything on a confidence score.

Novel Cognition made two errors drafting this analysis, both published at fallsover.novcog.us.com/corrections rather than corrected silently. One of them was dismissing a real technical term as a typo on sight, without checking the documentation — the same failure the analysis documents in others.

Prior analysis of what TypeSafe shipped, published the day after launch, is at jev.novcog.us.com.


Part of the Field Analysis Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(1523) "
Screenshot from TypeSafe evals explaining that reference labels are generated by GPT-6 Astra and Claude Fable 5.1.
A 33-minute explainer calls TypeSafe's Jev classifier one of the biggest events in the history of computer science. Its benchmark defines the correct answer as the average of two frontier LLMs, so it measures agreement rather than accuracy — TypeSafe says so itself. Every published speed measurement lands between 1.16x and 6x against a 193.6x headline, the lowest coming from TypeSafe's own employee. And a limitations section promised three times is absent from the video's own chapter markers." ["atom_content"]=> string(10584) "
Screenshot from TypeSafe evals explaining that reference labels are generated by GPT-6 Astra and Claude Fable 5.1.

By Bill Moran

TypeSafe is positioning its new Jev classifier as a disruptive force in computer science, promising a combination of high speed, low cost, and rigorous type guarantees. However, an analysis of the vendor’s claims and a recent 33-minute explainer suggests the actual economic impact may be overstated. While the pricing model is legitimate, the scale of the promised efficiencies and several key promotional claims are not supported by the source data.

What independent testing found

An external benchmark published 17 September ran 2,000 synthetic phishing emails, reproducibly, with no vendor relationship. Asked the straightforward question once — is this phishing — Jev scored 62.6%. Claude Haiku 4.5, a small and inexpensive model that has been shipping for a year, scored 81.3%.

A 95% figure from the same benchmark also circulates. Reaching it required splitting the single question into five narrower ones, hand-labelling 1,000 examples and fitting a logistic regression over the results. As the researchers put it, that 95% is not Jev; it is Jev plus your labelled data plus a regression you maintain — approximately the work a general-purpose classifier was sold to remove.

The most consequential result for anyone planning to route on Jev’s confidence scores came from a pre-registered evaluation on 20 September. Thirty messages falling outside the scope of every available answer were put in front of the model. It flagged none of them, at 0.99 confidence. On a separate unsolvable task it was correct 44.7% of the time while reporting an average probability of 0.74. Out of distribution, its expected calibration error measured 0.107, some 4.4 times the noise floor.

The type-safety guarantee is real but narrow: output cannot fall outside the schema. A forced-choice model cannot answer “none of these” unless that option was supplied, so an out-of-scope input does not produce an error. It produces a confident wrong answer. The model does not invent text. It invents certainty.

The zero is a definition, not a measurement.

Every published measurement lands between 1.16x and 6x

TypeSafe’s home page claims 193.6x faster and 444.6x cheaper. Set against that, the measurements anyone has actually published:

Both the large and the small numbers can be honest at once. The 193.6x figure compares a single decision call against a single frontier-model call on workflows the vendor wrote. The 15.9% figure measures what happens to an entire pipeline when one step inside it is swapped. Replace one component with something four hundred times cheaper and the system does not become four hundred times cheaper; the remaining calls still dominate.

Which is the point. The figure that matters to anyone deciding whether to adopt is the pipeline figure, and the pipeline figure is the one nobody quotes. The video’s presenter cites the most favourable real measurement available — six times — and then endorses a claim of roughly a hundred.

The full walkthrough, nine minutes:

Watch on YouTube, or read the full analysis at fallsover.novcog.us.com.

The benchmark defines the correct answer as what two other models say

Every speed and cost figure in the launch comes from an evaluation TypeSafe built. The company describes the scoring itself, on its own blog: “We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models.”

The consequence is structural. A perfect score on that benchmark would mean agreeing with those two frontier models on every question, including the ones they get wrong. Jev cannot outperform them, because they are the grading key. The number does not measure whether the model is correct about the world; it measures how closely and how cheaply it imitates the models it is being sold as a replacement for.

TypeSafe also records that the workflows “were made by individuals on our model capabilities team, so some bias could exist,” and that competing models were run through TypeSafe’s own wrapper. None of this is hidden. It is simply not carried along when the headline figure travels.

A related figure has the same shape. TypeSafe’s launch chart shows a 0% hallucination rate. The footnote beneath it reads: “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.” The zero is a definition, not a measurement.

A limitations section, promised three times

The video undertakes to explain where Jev falls over on three separate occasions: in the opening line of its description, again in the bulleted contents of that description, and twice in the narration, at 00:40 and 10:59.

What arrives, at 23:00, runs about twenty-five seconds: “Now, is Jev perfect? Of course not. It’s an excellent general purpose classifier. Like anything in computer science, there’s no perfect tool for everything. It still makes some mistakes.” No failure mode is named. No figure is given. Nothing in it is falsifiable.

This does not rest on a characterisation of the video’s structure, because its author published that structure. The chapter markers in the description read: 23:06 Testing Jev and getting started; 26:29 Cost and speed; 27:20 Jevons paradox and the future of software. There is no limitations chapter, because there is no limitations section. The twenty-five seconds above sit inside “Testing Jev and getting started.”

Material for such a section existed by 19 September. A reproducible outside benchmark had been published two days earlier. Calibration problems had been measured. None of it appears.

What survives

Three claims hold up without qualification. Because Jev answers in a single parallel pass rather than generating tokens one at a time, end-to-end latency of 70 to 500 milliseconds follows from the architecture. Because output is constrained to a caller-defined schema, malformed output cannot occur. And the pricing — $0.042 per million input tokens, with output unmetered — recomputes exactly; Novel Cognition checked the arithmetic on the assumption it was wrong, and it is not.

A fast, inexpensive, schema-locked classifier is a useful component, and the economics are worth testing. Three things are worth doing before relying on one: run it against your own labelled data rather than the vendor’s workflows; feed it inputs matching none of your options and observe what comes back; and measure calibration on your own distribution before routing anything on a confidence score.

Novel Cognition made two errors drafting this analysis, both published at fallsover.novcog.us.com/corrections rather than corrected silently. One of them was dismissing a real technical term as a typo on sight, without checking the documentation — the same failure the analysis documents in others.

Prior analysis of what TypeSafe shipped, published the day after launch, is at jev.novcog.us.com.


Part of the Field Analysis Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1790021136) } [9]=> array(11) { ["title"]=> string(124) "The Paper Everyone Cites for AI Search Tested Its Own Advice on a Real Engine. The Tactic Lost Three-Quarters of Its Effect." ["link"]=> string(111) "https://nocarolinachronicle.com/ai-search-optimization-five-assumptions-checked-primary-sources-north_carolina/" ["dc"]=> array(1) { ["creator"]=> string(10) "Bill Moran" } ["pubdate"]=> string(31) "Mon, 21 Sep 2026 04:04:29 +0000" ["category"]=> string(164) "NewsAI OverviewsAI searchcontent strategyCrowdStrikedigital marketinggenerative engine optimizationGEOGoogle patentinformation gain patentLLM citationsPerplexitySEO" ["guid"]=> string(40) "https://nocarolinachronicle.com/?p=53027" ["description"]=> string(1616) "
The Paper Everyone Cites for AI Search Tested Its Own Advice on a Real Engine. The Tactic Lost Three-Quarters of Its Effect.
Five load-bearing assumptions in AI search optimization, checked against the primary sources. Google's 'information gain' patent scores documents against what one reader already viewed, not site quality. The Princeton GEO paper validated its own tactics on a deployed engine in Section 6 — where adding statistics gains 8.7%, not the 30-40% headline. And one widely-shared workflow's showcase citation resolved to a vendor's marketing PDF sitting in a state police WordPress uploads folder." ["content"]=> array(1) { ["encoded"]=> string(10396) "
The Paper Everyone Cites for AI Search Tested Its Own Advice on a Real Engine. The Tactic Lost Three-Quarters of Its Effect.

By Bill Moran

An SEO newsletter recently published a Claude prompt that finds rare statistics buried in government PDFs. The prompt works: filetype-scoped retrieval genuinely surfaces material that page-one crawling misses.

The argument wrapped around it was wrong in five specific ways — and not one of those errors was original. Each is a load-bearing assumption underneath most of what the industry currently calls AI search optimization.

Novel Cognition checked all five against the primary sources: the patent, the paper, and the citation itself. Two of the five came back different from the first draft of this analysis, and those corrections are published rather than absorbed.

The patent does not say what the field thinks it says

US 11,354,342 B2, “Contextual estimation of link information gain,” filed by Victor Carbune and Pedro Gonnet Anders with a priority date of 18 October 2018 and assigned to Google LLC. Its abstract:

An information gain score for a given document is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user.

Previously viewed. By the user. This is a per-user, per-session, next-document score — the patent describes a first set of documents already presented to someone, then scores new documents on whether they contain anything that set didn’t. Its worked example is an automated assistant deciding what to read out loud next.

It is not a site quality score, and it does not say novel pages outrank derivative ones. It says: do not repeat to this specific reader what this specific reader just consumed. Anything done to game it is therefore conditional on a user state a publisher cannot see, measure or influence.

The steelman deserves stating. Google has kept prosecuting the invention: three grants are sourceable off that single 2018 filing, through US 12,013,887 B2 (granted 18 June 2024) and US 12,326,889 B2 (2025). Nobody prosecutes a claim family for seven years over something they shelved.

Five load-bearing assumptions in AI search optimization, checked against the primary sources. Google’s ‘information gain’ patent scores documents against what one reader already viewed, not site quali

The full walkthrough, twelve minutes:

Watch on YouTube, or read the full analysis at fivelies.novcog.us.com.

The paper tested its own advice on a real engine

The field’s single piece of evidence for “add statistics and get cited by AI” is GEO: Generative Engine Optimization (arXiv 2311.09735, six authors, accepted to KDD 2024). Its headline is a 30–40% relative improvement on a metric called Position-Adjusted Word Count.

What is almost never quoted is Section 6, where the authors ran the same tactics against Perplexity.ai — in their words, “a real deployed Generative Engine with a large user base.” Table 5, absolute values:

Method Position-Adjusted Word Count vs baseline Subjective Impression
No optimization 24.1 — 24.7
Keyword Stuffing 21.9 −9.1% 28.1
Quotation Addition 29.1 +20.7% 32.1
Statistics Addition 26.2 +8.7% 33.9

Three things follow. The tactic the entire pitch rests on gains 8.7% on a production system, not 30–40%. The 37% figure a marketer would quote comes from the other column — Subjective Impression is a language model’s judgement of how prominent content felt, which is neither visibility nor traffic. And Quotation Addition beat it better than two to one on the metric that counts words.

One more row worth keeping: keyword stuffing scored 9% worse than doing nothing. Traditional SEO did not merely fail to transfer. On this measurement it went backwards.

Four failures stacked in one citation

The showcase statistic in the workflow’s own demonstration: 79% of detections in 2024 were malware-free, sourced to the CrowdStrike Global Threat Report. The URL it resolved to:

fusion.vsp.virginia.gov/wp-content/uploads/

A Virginia State Police fusion centre’s WordPress uploads folder, hosting a copy of a private vendor’s report.

  1. inurl:gov selects for hosting, not authorship. The prompt instructs the model to prioritise government sources, and then launders a cybersecurity vendor’s own detection telemetry into a .gov authority signal. That is what the heuristic does structurally, every run.
  2. The URL is non-canonical. Citing a state agency’s upload directory rather than the publisher guarantees link rot and breaks entity resolution for any system trying to attribute the claim.
  3. The figure is vendor telemetry, not primary research — published by a company that sells the product the number argues for. That does not make it false; it makes it interested, and a citation without that flag has discarded what a reader most needs.
  4. The claim degraded in transit. CrowdStrike wrote “79% of the detections CrowdStrike observed were malware-free.” By the time it reaches a reader it says “79% of attacks.” One vendor’s sensor data became a fact about the world in a single hop, with nobody lying.

The figure has since been superseded: CrowdStrike’s 2026 report puts it at 82% of detections in 2025. The number moves. The relayed version does not.

What survives

Keep the retrieval — filetype-scoped search really does surface material page-one crawling misses. Use it as input discovery, then change everything downstream.

  1. Split the run in two. One phase retrieves candidates, a second verifies them. Never one continuous pass into a sealed document. One widely-shared prompt contains the instruction “Do not show the user a preview” — the single moment a human could catch an error, engineered out because previews make a demo look slow.
  2. Resolve to canonical source, always. Every statistic carries the publishing entity rather than the hosting domain, with a canonical URL, an original publication date, one line of methodology, and a commercial-interest flag.
  3. Verification gates the copy. Every number confirmed against the document itself, not against the extraction.
  4. Flip the value layer. Third-party statistics are context, never the differentiator. The gain is what an organisation can originate: its own instrumentation, measurement and distribution producing observations nobody else is positioned to make.
  5. Structure for the claim, not the page. A statistic buried mid-paragraph is not a retrievable unit.

The underlying shift: getting found is cheap now — three hundred sources in nine minutes. The question is what resolves to you.


Part of the Field Analysis Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" } ["summary"]=> string(1616) "
The Paper Everyone Cites for AI Search Tested Its Own Advice on a Real Engine. The Tactic Lost Three-Quarters of Its Effect.
Five load-bearing assumptions in AI search optimization, checked against the primary sources. Google's 'information gain' patent scores documents against what one reader already viewed, not site quality. The Princeton GEO paper validated its own tactics on a deployed engine in Section 6 — where adding statistics gains 8.7%, not the 30-40% headline. And one widely-shared workflow's showcase citation resolved to a vendor's marketing PDF sitting in a state police WordPress uploads folder." ["atom_content"]=> string(10396) "
The Paper Everyone Cites for AI Search Tested Its Own Advice on a Real Engine. The Tactic Lost Three-Quarters of Its Effect.

By Bill Moran

An SEO newsletter recently published a Claude prompt that finds rare statistics buried in government PDFs. The prompt works: filetype-scoped retrieval genuinely surfaces material that page-one crawling misses.

The argument wrapped around it was wrong in five specific ways — and not one of those errors was original. Each is a load-bearing assumption underneath most of what the industry currently calls AI search optimization.

Novel Cognition checked all five against the primary sources: the patent, the paper, and the citation itself. Two of the five came back different from the first draft of this analysis, and those corrections are published rather than absorbed.

The patent does not say what the field thinks it says

US 11,354,342 B2, “Contextual estimation of link information gain,” filed by Victor Carbune and Pedro Gonnet Anders with a priority date of 18 October 2018 and assigned to Google LLC. Its abstract:

An information gain score for a given document is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user.

Previously viewed. By the user. This is a per-user, per-session, next-document score — the patent describes a first set of documents already presented to someone, then scores new documents on whether they contain anything that set didn’t. Its worked example is an automated assistant deciding what to read out loud next.

It is not a site quality score, and it does not say novel pages outrank derivative ones. It says: do not repeat to this specific reader what this specific reader just consumed. Anything done to game it is therefore conditional on a user state a publisher cannot see, measure or influence.

The steelman deserves stating. Google has kept prosecuting the invention: three grants are sourceable off that single 2018 filing, through US 12,013,887 B2 (granted 18 June 2024) and US 12,326,889 B2 (2025). Nobody prosecutes a claim family for seven years over something they shelved.

Five load-bearing assumptions in AI search optimization, checked against the primary sources. Google’s ‘information gain’ patent scores documents against what one reader already viewed, not site quali

The full walkthrough, twelve minutes:

Watch on YouTube, or read the full analysis at fivelies.novcog.us.com.

The paper tested its own advice on a real engine

The field’s single piece of evidence for “add statistics and get cited by AI” is GEO: Generative Engine Optimization (arXiv 2311.09735, six authors, accepted to KDD 2024). Its headline is a 30–40% relative improvement on a metric called Position-Adjusted Word Count.

What is almost never quoted is Section 6, where the authors ran the same tactics against Perplexity.ai — in their words, “a real deployed Generative Engine with a large user base.” Table 5, absolute values:

Method Position-Adjusted Word Count vs baseline Subjective Impression
No optimization 24.1 — 24.7
Keyword Stuffing 21.9 −9.1% 28.1
Quotation Addition 29.1 +20.7% 32.1
Statistics Addition 26.2 +8.7% 33.9

Three things follow. The tactic the entire pitch rests on gains 8.7% on a production system, not 30–40%. The 37% figure a marketer would quote comes from the other column — Subjective Impression is a language model’s judgement of how prominent content felt, which is neither visibility nor traffic. And Quotation Addition beat it better than two to one on the metric that counts words.

One more row worth keeping: keyword stuffing scored 9% worse than doing nothing. Traditional SEO did not merely fail to transfer. On this measurement it went backwards.

Four failures stacked in one citation

The showcase statistic in the workflow’s own demonstration: 79% of detections in 2024 were malware-free, sourced to the CrowdStrike Global Threat Report. The URL it resolved to:

fusion.vsp.virginia.gov/wp-content/uploads/

A Virginia State Police fusion centre’s WordPress uploads folder, hosting a copy of a private vendor’s report.

  1. inurl:gov selects for hosting, not authorship. The prompt instructs the model to prioritise government sources, and then launders a cybersecurity vendor’s own detection telemetry into a .gov authority signal. That is what the heuristic does structurally, every run.
  2. The URL is non-canonical. Citing a state agency’s upload directory rather than the publisher guarantees link rot and breaks entity resolution for any system trying to attribute the claim.
  3. The figure is vendor telemetry, not primary research — published by a company that sells the product the number argues for. That does not make it false; it makes it interested, and a citation without that flag has discarded what a reader most needs.
  4. The claim degraded in transit. CrowdStrike wrote “79% of the detections CrowdStrike observed were malware-free.” By the time it reaches a reader it says “79% of attacks.” One vendor’s sensor data became a fact about the world in a single hop, with nobody lying.

The figure has since been superseded: CrowdStrike’s 2026 report puts it at 82% of detections in 2025. The number moves. The relayed version does not.

What survives

Keep the retrieval — filetype-scoped search really does surface material page-one crawling misses. Use it as input discovery, then change everything downstream.

  1. Split the run in two. One phase retrieves candidates, a second verifies them. Never one continuous pass into a sealed document. One widely-shared prompt contains the instruction “Do not show the user a preview” — the single moment a human could catch an error, engineered out because previews make a demo look slow.
  2. Resolve to canonical source, always. Every statistic carries the publishing entity rather than the hosting domain, with a canonical URL, an original publication date, one line of methodology, and a commercial-interest flag.
  3. Verification gates the copy. Every number confirmed against the document itself, not against the extraction.
  4. Flip the value layer. Third-party statistics are context, never the differentiator. The gain is what an organisation can originate: its own instrumentation, measurement and distribution producing observations nobody else is positioned to make.
  5. Structure for the claim, not the page. A statistic buried mid-paragraph is not a retrievable unit.

The underlying shift: getting found is cheap now — three hundred sources in nine minutes. The question is what resolves to you.


Part of the Field Analysis Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on Daily California Press
→ Coverage from FiorReports – Trump

" ["date_timestamp"]=> int(1789963469) } } ["channel"]=> array(8) { ["title"]=> string(24) "North Carolina Chronicle" ["link"]=> string(31) "https://nocarolinachronicle.com" ["description"]=> string(13) "Tarheel Times" ["lastbuilddate"]=> string(31) "Sat, 03 Oct 2026 19:50:11 +0000" ["language"]=> string(5) "en-US" ["sy"]=> array(2) { ["updateperiod"]=> string(9) " hourly " ["updatefrequency"]=> string(4) " 1 " } ["generator"]=> string(30) "https://wordpress.org/?v=7.1.2" ["tagline"]=> string(13) "Tarheel Times" } ["textinput"]=> array(0) { } ["image"]=> array(0) { } ["feed_type"]=> string(3) "RSS" ["feed_version"]=> string(3) "2.0" ["encoding"]=> string(5) "UTF-8" ["_source_encoding"]=> string(0) "" ["ERROR"]=> string(0) "" ["WARNING"]=> string(0) "" ["_CONTENT_CONSTRUCTS"]=> array(6) { [0]=> string(7) "content" [1]=> string(7) "summary" [2]=> string(4) "info" [3]=> string(5) "title" [4]=> string(7) "tagline" [5]=> string(9) "copyright" } ["_KNOWN_ENCODINGS"]=> array(3) { [0]=> string(5) "UTF-8" [1]=> string(8) "US-ASCII" [2]=> string(10) "ISO-8859-1" } ["stack"]=> array(0) { } ["inchannel"]=> bool(false) ["initem"]=> bool(false) ["incontent"]=> bool(false) ["intextinput"]=> bool(false) ["inimage"]=> bool(false) ["current_namespace"]=> bool(false) ["last_modified"]=> string(30) "Sat, 3 Oct 2026 22:23:10 GMT " }