Othello Merriweather, WordBlock Labs

I build local AI tools and test them honestly.

I'm an AI systems builder, not a coder. I design the workflows, the rules and the tests; AI tools write the code under my direction. What I'm good at is finding where a system fails, designing the fix, and proving the fix holds.

I started because I wanted a roleplay partner I could pick up for ten minutes, on my own computer, without anyone watching. So I built one, and then I measured it.

What I built

Three things, one way of working

SoloRoleplayer

A private, local, turn-based roleplay frontend in the style of the MUSH games I grew up on. You write your pose; the companion writes only its own. The app keeps the facts (where everyone is, what they wear, what was promised) and the model performs. No accounts, no telemetry, nothing leaves your computer. For adults, 18 and over.

Status: in testing; a Windows release is planned.

A trained companion model

A LoRA for Mistral Nemo, trained to keep a companion in its own lane (never writing for you) and to keep tense and quotation marks clean. Most of the training data was written by hand.

It will be published on Hugging Face (WordBlockLabs) with the release.

Michelangelo

My working method for AI-assisted projects: build in layers, test each layer, write down where you stopped. Breadcrumbs, a rolling ledger for every project and a control center keep a dozen projects organized and repeatable. It is also how I produced an original audio-drama series with scripted narration, sound design and a documented pipeline.

Evidence

I measure my own work, and I publish the weak spots

A first benchmark

Across four runs of 216 turns each on a pinned build, with the same seeds for both models, the LoRA cut head-hopping (writing the player's words or actions) in what the player sees from 24% of replies to 6%, and kept tense consistent in 99% of replies against 46% for the base model. In the model's own output, before the app touches it, the drop is smaller: 12% to 7%.

It also found problems: replies run short (47 words on average against 125), the longest setting met the paragraph minimum only 21% of the time, it repeats phrases more, and it broke one "address the player as Captain" rule 43% of the time against 9%. The testing also caught a bug in my own app's text cleanup, since fixed. Those problems are on the list to fix.

How it was run

Scripted player, deterministic scoring, a rented GPU, full proof folder with hashes, and a methodology page that lists the failed attempts and the known limits. It is a small four-run pilot; read it as a feasibility result, not a final claim.

The full methodology and results will be published with the release.

The story of how it was built

From a Canva sketch and a chat window to a tested frontend, in the order it happened, including what went wrong. The full write-up will be published with the release.

How I work

Stops, checks and repeatable results

Layers

Add one thing, test it, add the next. When I tried "just make X" I got a mess, so I built a process instead.

Stops and notes

Every stopping point gets written down so the next session starts from facts, not memory. Hand-off packets became breadcrumbs, and breadcrumbs became a ledger.

Plain about the tools

AI tools wrote most of the code. I designed the systems, decided what counts as a failure and tested every change. I say so up front.

Work with me

Looking for a role, or a collaborator?

I'm looking for work

Roles in AI companion and conversation design, AI evaluation and quality, content and workflow operations. Remote. I don't use social media, so this page is the best place to find me.

Résumé available on request: send a note through the support form and choose "General question or collaboration".

Questions and support

Questions about SoloRoleplayer or Cute LM, or a general note or collaboration idea: use the support form and pick the type that fits.

Open the form

Background

Before this: years in business-to-business account management at a Fortune 50 company, and three years running my own property-services company. I've also produced an original audio drama and a long run of explainer videos.