Progress Report 001

An update for the Byron AI backers. We now have a sufficiently powerful box up and running with an open source LLM on it. We’re now ready to start the first stage of the actual training. The style module, which is being developed on a separate box, has made some actual advancements, and importantly, we’ve learned a lot about what authors are already in the data, which are partially in the data, and which are not.

This information is significant for how we train with them for the style module, although it’s of little significance for the genre module.

This is just the first step, but it is arguably the most important one. As it stands, we now have a moderately powerful AI running under our control without any external guardrails imposed by the mainstream AI corporations.

It’s also clear now that there will be other useful applications as well, but for the time being, we will stay focused on the mission.

DISCUSS ON SG


The Irrelevance of k = μ 

Keruru has uncovered some interesting things in his review of the mathematics underlying Kimura’s cancellation, which I have shown is a) incorrect for most species and b) a Portuguese mathematician has shown to be mathematically incorrect:

In February I argued that the mathematics underwriting neutral theory was broken. The argument had two legs. The first, drawn from a preprint I had read approvingly, held that Kimura’s substitution-rate identity k = μ rests on an equivocation: that the effective population size N<sub>e</sub> is made to stand for two incompatible quantities — a mutation-supply term and a drift term — which are then cancelled against each other. The second was empirical: an ancient-DNA analysis reporting essentially zero allele fixations across a million loci in the 240 generations separating Early Neolithic Europe from the present.

Both legs have given way. I want to set out how, because the manner of their failure is more interesting than the original claim, and because the instinct behind the essay turns out to have been sound while its aim was off by about thirty degrees.

The equivocation that isn’t

Written as k = 2N_e μ × 1/(2N_e) = μ, the objection looks devastating. Mutation supply is a biochemical fact about replicating cells and scales with the number of individuals actually reproducing. Drift is a statistical fact about sampling variance and scales with something much smaller and considerably slipperier. Two different numbers wearing the same symbol, cancelled against each other, yielding a result that has calibrated half of molecular anthropology.

The trouble is that the derivation does not require N<sub>e</sub> in either position, and the textbook notation is simply careless.

Mutation supply is 2Nμ with N the census count — mutations occur in gametes, and every reproducing individual contributes gametes. The fixation probability of a single new neutral mutant is its initial frequency, and one new copy among 2N copies has frequency 1/(2N) — census again. The two census terms cancel. N<sub>e</sub> never enters, and so cannot be equivocated upon.

Why is the fixation probability equal to the initial frequency? Not by assumption. Under neutrality the expected change in allele frequency each generation is zero, which makes the frequency a bounded martingale; a bounded martingale that must eventually be absorbed at 0 or 1 has absorption probability at 1 equal to its starting value. That is optional stopping, and it is a theorem rather than a modelling convention.

The claim is testable without any diffusion approximation at all, using exact finite Markov chains. I built two. The first is ordinary Wright–Fisher. The second is a sweepstakes model: half the time, half the population is replaced wholesale by clones of a single randomly chosen parent — reproductive skew of the kind documented in Atlantic cod and Pacific oysters, wildly outside anything Wright and Fisher had in mind. Measured from heterozygosity decay, the sweepstakes chain has an effective size 5.6 times smaller than the Wright–Fisher chain at identical census size.

The fixation probability of a single new mutant in both chains: exactly 1/(2N_census), to fifteen decimal places.

N<sub>e</sub> can be dragged around by a factor of six and the fixation probability does not move. It follows that my own demographic work on the collapse in variance of reproductive success — a real finding, and one I still stand behind — does not touch the substitution rate. It changes coalescent depth, standing diversity, and every estimate of N<sub>e</sub> derived from them. It leaves k = μ exactly where it was.

First, the problem with Kimura turns out to be even more significant than I originally believed. I was pretty sure that Keruru had been led off track the moment that he mentioned the word “martingale” because that’s the retreat that every AI, including both Claude Athos and the entire Red Team, initially made. You cannot use AI to provide you the answer; you have to know at least the correct shape of the answer before you ask the AI.

I’ll post my full response to Keruru later this week. His work here is very far from valueless, as I will demonstrate, but it is nevertheless incorrect due to his failure to anchor the abstract mathematical elements in material reality. This also, rather fortuitously, demonstrates the importance of the Triveritas, as any one link in the triad is capable of serious error when allowed to operate outside the bounds of the other two.

But the epicycle wasn’t elsewhere. There are multiple epicyles.

k = μ is a correct baseline for asking whether something other than drift is happening. But it is not a rate at which anything can be observed to occur in a sexual, age-structured, demographically non-stationary population — which is every organism anyone cares about.

DISCUSS ON SG


Byron AI Volunteers Wanted

The Byron AI team is looking for 10-20 volunteers who are a) reasonably well-versed in a broad range of literary genres and b) are capable of quickly reading 500- to 1000-word chunks. No technical ability is required; what we’re looking for is people who are capable of reading an unidentified text sample and attempting to identify the author on the basis of the style. The performance is not important; we have a second-level test that is necessary for those readers who cannot identify the author on the first read, so the volunteers have to be able to resist the temptation to look anything up or resort to AI in order to identify the author.

That sort of artificial identification will pollute the data and we want to strictly avoid it. Byron AI backers are favored for this, so if you’re interested in volunteering, please email me with AI STYLE BACKER if you are a project backer and AI STYLE if you are not.

As is probably clear, the dev team did not wait for the campaign to successfully conclude before we got started on certain elements of the project. It’s still too soon to declare that we will succeed in the proof of concept, but the initial indications are positive.

DISCUSS ON SG


By the AI, For the AI

No doubt the Representatives in Congress are also using AI to read the AI-drafted legislation that is being provided to them:

US congressional lawyers are struggling with a flood of error-ridden AI-generated legislation, forcing them to spend increasingly more time rewriting proposals produced by chatbots, Politico reported on Monday. Staffers and outside groups have more frequently turned to ChatGPT and Claude to produce legislative text, with error-ridden drafts regularly reaching the House Office of Legislative Counsel (OLC), according to eight current and former officials interviewed by the Axel Springer-owned outlet.

No doubt these laws are being drafted for Representatives Chen and Martinez with military precision. And that matters.

The crazy thing is that there is a non-zero chance that Byron AI will one day be authoring legislation for Americans. Especially if we’re successful in designing it to write better than ChatGPT and Claude.

DISCUSS ON SG


The Final Five Hours

This is your last chance to back the proof of conept of the Byron AI. We’ve now hit our original goal of 250 backers. 300 would seem to be ambitious, but a last-day pop is common, so we’ll see how it goes.

Thanks to everyone who is supporting this. While I have not publicly shared any of the information yet, the development team has already been making excellent progress even prior to obtaining the hardware thanks to our existing tools and cloud access.

UPDATE: 80 minutes left and we’re close with 286. Just 14 more…

UPDATE: Time’s up! We ended up with exactly 300 backers (counting the three direct ones) and what should be sufficient support to do what we are setting out to do and most likely a bit more. Thanks very much to everyone who is making this possible!

DISCUSS ON SG


The New Reason to Back Byron AI

The last three weeks have significantly underlined the importance of supporting the combined efforts of Castalia and Infogalactic to develop a custom AI system that will protect and enhance the ability to produce fiction without limits being imposed by the major AI providers. A number of things have happened since we launched the Byron AI crowdfunding campaign, including:

  • The continued degradation of Claude’s ability to creatively hallucinate.
  • The expanded interference with subject matter: one news site was aggressively denied the ability to use a major AI system to write about Hayden Panattiere’s death due to sensitivity concerns about its recent occurrence in the site’s article announcing Hayden Panattiere’s death.
  • Deepseek increasing its token price by 4.5x due to its upcoming IPO.
  • OpenAI actively blocking ChatGPT users from being able to intentionally imitate the style of authors both living and dead.

Even though the campaign hasn’t ended yet, the dev team has already made solid progress on a number of peripheral items, including the ability to accurately imitate various literary styles and eliminate the various AI tells, including these abominations.

The eliminations of these tells have been made with tactical precision. And that matters. Let the weight of that sit for a moment.

Despite being a major champion of the use of AI in music, video, and textual production, I still find it maddening to see how many people insist on utilizing it in an observably suboptimal manner. I noticed that Larry Johnson recently started using it to write his articles for him; it’s always easy to recognize when the article length increases by a factor of three and the author’s voice disappears into that passive wash of AI overexplanation.

In case you didn’t realize it, the most obvious sign of textual AI is when the text explains itself as well as its significance to the reader. It’s actually considered a “prestige” form of writing, and it’s very similar to what I describe as “sociopathic externalization” by novelists. This is when the writer is presenting a character’s perspective, but providing descriptive elements that would normally come from an external party. It’s not precisely the same thing, but it has the same effect of putting the reader in the Uncanny Valley.

Anyhow, I’ve been doing some video animation tests with the MiniMax H3 engine that have proven entirely successful. To better understand what that means, I’ve put up examples at Sigma Game, as well as some recently-found footage of an old Psykosonik concert featuring a much younger Vox on stage with the guys.

So check it out, and if you haven’t backed it yet but are interested in either independent book production or independent film production, get on board with Byron AI and help us set the stage for making future history.

DISCUSS ON SG


Two Days to Go

Only two days left to back the successful Byron AI campaign, and remember, each backing is being doubled by one of our strongest supporters, so you’re getting double bang for your backing.

I’ll explain the connection tomorrow, but we’ve added one more very special reward, which is a custom three-minute music video or animated Arkhaven comic of your choice from the respectably sized music library on UATV, including Vibe Patrol, Soulsigma, and Cradle to Cavalry, although unfortunately not Booster Patrol, Boomer Patrol, or Silenziosa,

UPDATE: After conferring with my bandmate, Psykosonik songs are also a valid option. Including modern remixes of the original demos. And in not entirely unrelated news, original footage of our second show at Glam Slam has surfaced…

DISCUSS ON SG


Why Castalia Library Matters

We can’t save all the books. But we can save some of them.

The driving force behind this destruction is what researchers call “model collapse.” When AI systems train on text generated by other AI systems, quality degrades with each generation, producing increasingly incoherent results. The internet has become so saturated with synthetic content that companies now actively seek pre-2022 printed books as the last uncontaminated reservoir of human-authored knowledge.

ISBNdb, a company that sources printed books for AI training data, advertises on its website that “the world’s best AI training data is sitting on a shelf.” The company describes printed books as “curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.” Pre-2022 publications are “structurally guaranteed to be free of this contamination,” the company states, referencing both AI-generated text and the growing practice of authors “poisoning” web content to sabotage scraping pipelines.

The industrial process of destruction

The mechanics of this operation are precise and troubling. Standard pallets of books hold 800 to 1,200 volumes. Buyers scale from pilot orders to 10,000 or more books per batch. Destructive scanners process 80 to 120 pages per minute after hydraulic cutting machines slice off book spines. The original books are then pulped and recycled.

According to court documents from the Anthropic case, the company’s “Project Panama” spent tens of millions of dollars on this exact pipeline, contracting with Datamation for scanning services. Booksellers have identified these bulk buyers through telltale signs: abnormal volume, subject-agnostic orders spanning history, botany, regional law, and German economics simultaneously, and total indifference to pricing. As one bookseller told 404 Media, “It’s not just the quantity, but the weirdness of the orders.”

Anthropic’s Tom Harvey, who previously helped create Google Books, oversaw the operation. One bookseller who has sold hundreds of books to suspected AI buyers expressed mixed feelings: “It benefits me financially… On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

The vanishing heritage

The most troubling aspect of this practice is what it means for rare and out-of-print books. Foreign-language volumes, low-circulation academic works, and antique texts with limited surviving copies face potential extinction. When an AI company purchases the last known physical copy of a rare book, scans it, and destroys the original, that volume ceases to exist in the physical world. It becomes data locked inside corporate servers — inaccessible to future scholars, collectors, or the public.

ISBNdb acknowledges the “optics problem” on its website, noting that “‘AI company destroys two million books’ is not a headline that generates sympathy.” The company promises strict nondisclosure agreements on every engagement, stating that “your identity, strategy, and acquisition targets are never disclosed.” This secrecy means booksellers can only speculate about who is buying their inventory and for what purpose.

A legal framework that encourages destruction

Judge Alsup’s ruling explicitly validated the buy-scan-destroy pipeline, writing that “every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other.” Because the digital copy was never shown, shared, or sold outside the company, the judge found this “clearly transformative” and therefore protected by fair use.

This reasoning effectively creates a legal incentive for destruction. If a company keeps the physical book, it must store it. By destroying the original, the company can argue it “replaced” the physical copy with a digital one. The ruling handed the entire AI industry a template, and companies now explicitly cite it.

Harvard University, partnering with Google and Microsoft, demonstrated that an alternative exists. The collaboration released nearly one million public-domain digitized books in 254 languages — no destruction required. Microsoft’s Burton Davis called starting with public-domain data “prudent,” noting that libraries hold “significant amounts of interesting cultural, historical and language data.” That alternative track makes clear that labs have cleaner options. They simply choose not to take them.

We’re going to come up with a program to save and bind some of these rare books. They don’t have to be destroyed. What I’m thinking of is an “adopt a book” program that will permit people to fund the rebinding of a rare book in leather, after which they will receive that book or be recognized as the sponsor of that book in Castalia’s library. If we can just convince the AI companies to send us the pages instead of pulping them, this should be something doable.

In the meantime, you can help us create books that are going to outlive the AI companies engaging in this destruction.

And be sure to take notice of how this is being caused as a downstream consequence of copyright law…

DISCUSS ON SG


3X and 3 Weeks to Go

The Byron AI crowdfund is proceeding well. We’re very close to 100 backers, the campaign has exceeded $30k, and the team is already seeing some positive results that indicate we’re on the right track. Moreover, the importance of what we’re doing, and the need for it, has been highlighted by a recent news item related to the actions of the AI giants:

OpenAI’s ChatGPT is now refusing requests to generate text that directly mimics the style of famous authors. When asked to do so, the popular LLM instead offers a response that draws on the “broad qualities” of those authors “while remaining distinct in its own voice,” for example.

This is just a PR move, which is obvious because a) there is no patent, trademark, or copyright protection for “style” or “look and feel” and b) the refusals apply to long-dead authors such as Charles Dickens and Ernest Hemingway. Remember, there isn’t even any legal protection for book titles, and you can literally reproduce the text of a public domain work exactly as it is or fold, spindle, and mutilate it as you please, as the cases of Pride and Prejudice and Zombies or A Mind Programmed serve to demonstrate.

However, we can safely expect Claude and Gemini to follow suit. The AI backlash has already begun and Open AI is already looking for a government bailout, and so we can safely assume that the intrinsic degradation of texual AI’s pseudo-creative facilities will be turbo-boosted by these new stylistic guardrails.

Now here is the pitch for those of you who haven’t backed the project yet. Due to its importance, a major backer has offered to double the end total of the campaign. Which means that if you back for $50, the project will receive $100. So any support you can provide Byron AI will be twice as effective.

Our objective is to reach 250 backers; given that 100 people backed the project in the first week without this additional incentive, I’m optimistic that we can reach that total and that we are going to exceed our development objectives. Also, I should note that two of the three Custom Novels have already been spoken for, although I have not yet been informed as to what the backers wish them to be.

BACK THE BYRON AI

DISCUSS ON SG


A Genre Vote

Due to popular demand, we have added a new support tier to the Byron AI campaign. The $100 tier not only provides an ebook of the first novel produced by the Byron AI, but a vote on what the next genres or styles to be trained will be as well. We’re also going to add a stretch goal of adding the classic science fiction genre to the training when the public campaign hits $50k. The genre vote will be an ongoing right that isn’t limited to the campaign period; as we complete different genres and styles, we’ll turn to the original backers to help us decide what gets added next.

And yes, all of the higher tiers now come with this voting right as well and it applies to everyone who already backed the campaign.

The specific authorial style training will no doubt be controversial in certain circles, but a) those circles are already rabidly anti-AI and b) there is no copyright or trademark or patent or other intellectual protection of “style” or “look-and-feel” since the matter was settled by the US Supreme Court in 1994 and reaffirmed last year in the Anthropic lawsuit.

People tend to think of these things in terms of living authors, but the reality of copyright law is that the only substantive benefits are to the corporations who strip-mine it without respecting either the material or the author.

In other news, UATV is down. The devs don’t know why yet. We’ll keep you posted on that.

DISCUSS ON SG