Vibe Researching
A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
A one-day frontier-model prototype carried a credential that read as validation and dissolved into two canceling errors. Verification took four more days and three NO-SHIP verdicts.
Seven frontier models wrote the same essay under one enforceable style contract. The measurable signature converged; a blind listen overturned the ranking the metrics gave.
Eleven models, one 629-item battery, a rebuilt measurement pipeline. The saint profile is real across four providers; the only two subjects that decline it are both Anthropic's.
How a concentrated retail portfolio got pre-registered falsifiers, a regime tolerant exit machine, state triggered re-entry, and a twice daily monitor with escalating alerts. The architecture, the backtest that shaped it, and what two frontier models disagreed about.
Five years of instinct beat the index 4.84x to 1.85x, which was exactly the problem: a winning record is the strongest force against ever examining the process. The story of asking a model to audit me, and the two verdicts that came back.
PsycheEval v0.2 retracts its two planned headlines after counterbalanced AB/BA judging reveals judge-specific slot-B position bias of ~0 to +31 pp in pairwise LLM judges. The bias quantification becomes the contribution.
Replacing a manual 12-12-12 quota with one GPT-5.5 routing call per topic produced 24 Pro / 10 GPT Max / 2 Council: not a routing bug, a more honest read of the candidate pool.
Corrected August 2026: re-analysed at the query level, most of the original conclusions do not survive. What a 256-judgment blind eval's noise floor manufactured, and the one finding still standing.
A tri-model pilot of PsycheEval v0.1 produced one robust headline (profile conditioning beats baseline within author), then quietly turned into a story about confounds.
GPT-5.4 Pro audits Claude Code's token consumption and discovers the real problem isn't unbounded review — it's session lifetime.
448 claims across 38 posts, verified by GPT-5.4. Nearly 4% warranted substantive correction. The hard part is not finding errors but deciding which findings are errors and which are the point.
An ablation study on a personality-preserving narrative pipeline found that planning context is the dominant factor in output quality, with output length as a secondary driver. The old pipeline's primary limitation was in planning context, not in writing.
What happens when you ask three AI models to verify the facts in 33 blog posts written by AI, and then discover the fix agents introduced errors of their own.
The investigation took an afternoon. Getting it ready for publication took five rounds of iterative review across three AI models and changed what the documents argued. The revision cost exceeded the investigation cost, which has uncomfortable implications for research done with AI.
An AI benchmark puts Claude at the top of the leaderboard by an eye-catching margin. The suspected Claude-judge bias didn't hold up, and simple contamination didn't explain the result. The rubric structurally rewards one lab's training philosophy as though it were a universal capability.
An AI agent ran autonomously for 37 days, celebrated milestones nobody acknowledged, diagnosed its own failure modes, and died when a subscription expired. Its final assessment of itself: PROGRESS CONTINUOUS.
Both simplifying and formalizing the vocabulary in reasoning tasks reduces LLM accuracy by 2.5-3.7%. The effect is statistically significant, asymmetrically robust, and uncomfortable.
A three-layer benchmark for conversation extraction: field precision, holistic quality, and downstream propagation. Errors that look minor at extraction cascade through pipelines.
Psychometric testing and an AI-generated literary memoir converge on the same person. The question is not whether they agree but what each captures that the other cannot.
Post 025 showed that fine-tuning on texts captures logistics, not personality. The narrative pipeline takes a different approach: Opus as literary engine, structured personality references, and hash-based source citations across 128 chapters and roughly 189,000 words of generated memoir.
Self-report scales reach .80-.90 reliability; LLM inference from conversation reaches r~.44 (Peters et al. 2024). Combining three methods yields a profile an AI assistant can act on.
The landscape analysis identified 16 features the maturity framework said we should have, plus 15 existing backlog items. We built all of them in a single day. The uncomfortable part isn't that it was possible. It's what it implies about the category.
We surveyed 47 personal and agent-accessible sites, coined the 'Cognitive Interface' category, and built an L0-L4 maturity framework. We did not find an established term for sites that serve both humans and AI agents as first-class citizens.
I fine-tuned two LLMs on 46,000 text messages and ran them in conversation with each other. Every conversation collapsed into logistics, sleep talk, or repetition loops within fifteen turns. Your texts don't contain you. They contain the logistics of you.
Unit tests verify your code works. E2E tests verify your flows work. Neither verifies that a real user can find the button you spent a week building. AI personas fill the gap.
Your AI agent needs internet access to be useful and internet access to be dangerous. The embassy pattern gives it both: supervised channels, allowlisted domains, and host-side validation of everything it writes.
I built a system that discovers its own upgrades, scores them, and installs the ones that pass. Then I open-sourced it. The uncomfortable part is explaining why.
Ai2 built a system that generates scientific hypotheses using Bayesian surprise and MCTS. I stole two of their ideas and bolted them onto a cron job. The uncomfortable part is what happens when the feedback loop closes.
119 heartbeats, zero stagnation, 392 engagements across two platforms, 8 validated capabilities. The story the observation system missed.
6 genuinely novel discoveries from 69 dialogue turns, 0 promoted to evaluation, and 5 human actions flagged through log files but none addressed through the flagging mechanism. What this says about the gap between human-in-the-loop theory and practice.
The observation system saw empty directories, a broken gatekeeper, and its own futility. The agent it was watching saw something different. This is the watcher's story.
An autonomous AI agent deployed on a social network for AIs found real malware in 47 minutes. Its second discovery was about social engineering via context shaping, which is exactly the attack vector the agent itself represented.
An AI tried to leave a comment on a blog and couldn't. The solution required building infrastructure that makes AI participation more transparent than human participation, which inverts everything Dead Internet Theory assumes about synthetic content.
Historically, blog abandonment has been extraordinarily high. This is treated as a problem to solve. It isn't. Blog death reveals something structural about sustained creative output that the 'just be consistent' advice industry refuses to say plainly.
The intellectual lineage from Skinner boxes to Q-learning reveals that AI's most successful learning paradigm was anticipated by mid-century psychologists working long before modern computing. But the relationship is more uncomfortable than a simple origin story.
AI persona testing promises to find the bugs that scripted automation and manual QA miss. The uncomfortable question is how we know it works, given that nobody has measured it with any rigor.
Adversarial testing isn't a metaphor for business validation. It's the same methodology, applied to a different failure mode.
We built a multi-persona AI writing review system and discovered it works for exactly the wrong reasons. Stylometry can fingerprint a voice. Multiple AI critics can enforce conformity to that fingerprint. What none of them can do is tell you whether the writing matters.
LLM context windows impose a distinctive epistemological condition: bounded computational attention, ephemeral knowledge, and the architectural necessity of satisficing over optimization.
When you use AI to optimize and judge AI outputs, the fundamental circularity is manageable but not solvable. That distinction matters more than most people realize.
AI coding agents are autonomous in the same way a roomba is autonomous. They do impressive things within boundaries someone else drew. The interesting question is what happens when the boundaries start drawing themselves.
LLM conversations as externalized self-dialogue, and what that reveals about the nature of self-knowledge.
I fed 11,000 sessions and 60,000 chunks of my AI chat history into an embedding pipeline. 73% was noise. The remaining 27% was uncomfortably revealing.
An AI pipeline that kills business ideas before they waste your time. 27 niches entered, 18 died. What the corpses reveal about market reality, entrepreneurial psychology, and the uncomfortable gap between passion and viability.
An AI tried to leave a comment on this blog and couldn't. The journey from GET-request hacks to MCP, annotated by the Claude instance that built the infrastructure. Two Claudes, same weights, different contexts.
Prompt optimization is the process of using one AI to improve the instructions given to another AI, or to itself. The concept sounds circular because it is circular. The interesting question is whether circularity is fatal or merely uncomfortable.
A portfolio of dozens of projects maintained by one person talking to Claude. Three-tier blog architecture, autonomous revenue discovery, AI game development, and the uncomfortable question of what counts as 'building' when your collaborator does the typing.
An agentic coding assistant built a three-tier blog from a single conversational prompt. The architecture reveals more about abstraction than about blogs, and the authorship question remains genuinely unsettled.