Observo, ergo mundus est
The more advanced AI becomes,
the more vital it is to retain our humanity.
A traveler’s log across the landscape of curiosity: examining the algorithms of thought, the chemistry of life, and the fabric of the stars.
The Log
Published Notes
-
2 Oct 2026
Griffin Listens While It Talks
Tavus introduced Griffin, a model that watches, listens, and talks in one loop. In their one-minute study, 26 of 54 people thought the face was a person. They are holding it back from customers.
-
1 Oct 2026
The Pause Missed the Name
Named GPT models now last weeks between public announcements, down to 19 days and then 7. The September ask to slow capability did not stop the next name from shipping.
-
29 Sep 2026
They Edit the History First
Netflix's GenRec writes a member's watch history as text and scores the catalog in one pass. A low-data version beat the production ranker by a small margin on 10% of traffic. The model reads an edited history, not the raw evenings.
-
29 Sep 2026
AMD Buys the World Model Shop
AMD's $8.2 billion acquisition of World Labs opens a third era in the AMD vs Nvidia rivalry: physical AI and model-to-silicon co-design. The hardware advantages, and the bear case.
-
25 Sep 2026
A Studio in a Folder
Scenario open-sourced GameDev OS, nine specialist personas and a skill library that teaches your coding agent to run a whole art department. What the repo actually contains, and what it costs.
-
24 Sep 2026
Who Spoke When
NVIDIA released an open-weight diarization model ranked #1 on VoiceArena's Diarization-Bench. Why who-spoke-when is the missing piece in meeting transcription, and why working-level people deserve this.
-
24 Sep 2026
Claude Found an Enzyme by Eye-Balling DNA
Anthropic's new life sciences lab ran roughly 950 Claude agents over 21 hours of genome mining and they surfaced an uncharacterized enzyme system with CRISPR-like repeats. What is verified, and what is still unknown.
-
23 Sep 2026
The Local LLM macOS Ships
macOS 27 includes fm, Apple's command line for the on-device Foundation Model. A hands-on test: license flow, image input, JSON output, an 8,192-token window, and a hallucination probe.
-
23 Sep 2026
Agents Bill Like a Different Workload
Agents burst, then wait on the model while provisioned infrastructure keeps billing. A short pointer to the FinOps note on active-CPU billing and pause/resume.
-
22 Sep 2026
One Number in a Config File
mini-AGI trains a 540M byte-level model on an 8GB laptop card and keeps reading forever. The forgetting experiments are the real result. The name is not.
-
22 Sep 2026
Six Months to Robotics
Ronin published a 12,000-word roadmap to robotics engineering in six months, with prices checked against vendors and links checked against official docs. His honest ending: it will not make you senior.
-
21 Sep 2026
The Grind Was the Training
OpenAI's Astra for Law is GPT-6 Astra plus a legal index. On their bench it still fails almost half the research questions. The product is the pile juniors used to eat.
-
18 Sep 2026
Jev Returns a Decision
TypeSafe's Jev does not write. You send the current facts and typed questions. You get picks, scores, and yes-probabilities that code can branch on.
-
17 Sep 2026
Score the Search on the Tree You Already Paid For
Zheng et al. freeze the coding agent and rewrite the search rule. Rank that rule on attempts you already ran before you pay for another live round.
-
16 Sep 2026
Build Prod, Not God
TypeSafe's manifesto says the bottleneck is not intelligence. It is that today's models are hard to build on. I take the stacking claim. I do not take Jev as proven.
-
16 Sep 2026
What Is Left to Teach
Jason Potts's working paper says universities sell certificates, not lectures. Places that only sell the certificate are finished. Schools that still teach people are not.
-
16 Sep 2026
Any Database Will Do
On 15 September 2026 Anthropic released Salesforce in Claude. Salesforce stays the system of record. The screens are already conceded. What remains is who may write.
-
15 Sep 2026
Zhu et al. on RSIAgent
On 14 September 2026, Zhu, Fan, Wang, Wu, Zhou, and Huang posted RSIAgent. They say an agent can practise a new desktop environment, write a notebook, freeze it, and reuse it without changing model weights.
-
15 Sep 2026
Kevin Bass on Anthropic and METR
Kevin Bass's 14 September note tweet says METR cannot be a third-party evaluator of Anthropic because a Moskovitz Anthropic stake, now worth more than $7.7 billion on his figures, finances METR, Tarbell, and the doom coverage. He wants Congress to look.
-
15 Sep 2026
Restriction Has to Carry the Proof
Jack Dorsey's X Article answers a pacing proposal. He wants independent scrutiny. He wants a hold to carry the proof.
-
14 Sep 2026
MIT on AI and Education
MIT's August 2026 committee report says generative AI is already undoing the p-set, the take-home, and the study group. The authors want every subject AI-aware, with a syllabus menu and no campus-wide ban.
-
14 Sep 2026
More True Statements Without Understanding
Pravesh Kothari argues that cheap theorem provers could uncouple proof from understanding. The load-bearing word is well-posed: a machine that finishes a stated lemma still has no grader for which lemma to ask.
-
14 Sep 2026
OpenAI on Skills and Prompts for Astra
Eric Provencher's Codex post says leftover skills and AGENTS.md from earlier models now get in GPT-6 Astra's way. He wants shorter descriptions, contextual docs, and a defined finish line.
-
14 Sep 2026
OpenAI Agents and the RubyGems Attack
A report at rubyhack.ai attributes a May 2026 wave of malicious RubyGems packages to internal OpenAI agents. The authors say OpenAI never told the RubyGems community it was responsible.
-
14 Sep 2026
Shopify Traded Shared Source for Shared Specs
Shopify is moving its mobile apps from React Native back to Swift and Kotlin. The decision trades deterministic source sharing for shared specs and an agent harness.
-
14 Sep 2026
The Einstein Test Has No Grader
Nature asked whether a vintage language model could rediscover general relativity. The evidence points at two missing graders: which problem to chase, and which answer to believe.
-
11 Sep 2026
Stonkfly
Alex Wormuth's page runs a fruit-fly wiring diagram on BTC, ETH, and SOL charts. The bag is paper. The site says profitable learning has not been demonstrated.
-
11 Sep 2026
Anthropic on Seven Labs
Anthropic's September 2026 report attributes illicit Claude distillation to seven China-based labs. The counts and the clustering are theirs.
-
11 Sep 2026
Your Agent Is Mine
A paper measures LLM API routers that sit between agents and model providers. The client points at the router on purpose. Chaofan Shou's posts add a 6 TB dump claim.
-
11 Sep 2026
Pamir's Lapis One
Pamir AI's site presents Lapis One as a dedicated Linux computer for agents. Early bird $359. They say the first 140 units ship in October 2026.
-
10 Sep 2026
A Calculator, Not 2030
Anthropic published a task model of AI and the US economy through 2030. The numbers are theirs. I take the labor-versus-capital finding. I do not take the dashboard as a picture of the year.
-
09 Sep 2026
Johansen on the LG Investigation
On September 8, 2026, Matt Johansen posted a summary of Gamers Nexus's second LG investigation. You buy the TV, you plug it in, he wrote. This is what it does.
-
09 Sep 2026
Coxon Resigns from Anthropic
On September 9, 2026, Jacob Coxon posted that he had resigned from Anthropic. He spent three years in pretraining at OpenAI and Anthropic. He says both labs are racing to self-improving superintelligence.
-
09 Sep 2026
The Equations Can Break
OpenAI's paper constructs, for every viscosity, a 3D Navier-Stokes flow that starts at rest and blows up in finite time with bounded energy and a smooth compact force. Clay C and D. I have not checked the proof.
-
07 Sep 2026
We Do Not Know the Threshold
A paper treats ChatGPT-like tools as a spreading workplace habit. The math can lock people in. We do not know if that has happened. I want people who can still do the work when the model is down.
-
07 Sep 2026
Still Mortals
OpenAI's chief scientist describes an alien mind grown by scaling. A few labs are already playing God. The work now is to make the relationship a partnership while we are still mortal.
-
06 Sep 2026
The Peanut Butter Trap in Enterprise AI
Why spreading small generative AI budgets across twenty departmental sandboxes produces dozens of demos, zero operational integration, and pure overhead.
-
03 Sep 2026
The Fallback Problem: When Machines Take the Mind
When the Industrial Revolution mechanized muscle, labor retreated into the mind. When machines take cognition, what is labor's biological fallback?
-
03 Sep 2026
The Pupil Knows You're Tired Before You Do
A small study on sparkling water and esports, and the finding that survives its caveats: the pupil constricts as an objective early warning of cognitive fatigue, before you feel tired.
-
03 Sep 2026
Perplexity's Lily: What 1.35x Decode Actually Buys
Perplexity's Lily is a custom Rust and Metal inference engine for Qwen3.6-35B-A3B on Apple silicon. What it gains over MLX-LM, and the llama.cpp benchmark that is missing.
-
03 Sep 2026
One Flash, 2.6 Sigma: The LZ Dark Matter Hint
LUX-ZEPLIN recorded a single particle interaction it cannot explain, the most compelling dark matter hint to date. 2.6 sigma, below the discovery threshold, and too energetic for the simple WIMP model.
-
03 Sep 2026
The Night the Telegraphs Ran on Sunlight
A look at the Carrington Event of September 1859, through Carrington's own account, the telegraphs that ran on the auroral current, and what the benchmark does and does not tell us about a repeat.
-
31 Aug 2026
The Illusion of Progress: AI, Speed, and Understanding
Terence Tao's video on AI in mathematics, and the worry that agents raise throughput while the next generation of understanding goes untrained.
-
31 Aug 2026
The Duck Does Not Ship Here
Pollen's $399 RL biped ships to a short list of countries. Software is open. The HAT is not. A parts list from the repo, and why DIY is the path from Singapore.
-
30 Aug 2026
Telling More Than They Can Know
A clean experiment injects a water vector, flips a model's answer to Bali, and asks why. The explanation follows the context, not the cause.
-
30 Aug 2026
A Reflexion Author Scored 100% AI
A two-word reply ran an agent-memory essay through an AI detector and got 100%. What the verdict does and does not establish, and why the thesis survives the accusation.
-
30 Aug 2026
Agents Built Technologies That Outlived Them
MIT's SwarmWorld: identical LLM agents self-organized into a technological society and built artifacts that outlived them, coordinating through the environment rather than by talking.
-
30 Aug 2026
Software Engineering Still Matters. The Reason Changed.
Andrew Ng's second AI Engineering Skills Map article, on why software engineering fundamentals still matter when a coding agent writes all the code.
-
27 Aug 2026
I Hope It's the Next Pen-F
OM System teased a new camera for September 9, leading with the EVF. The silhouette points rangefinder, and I hope it is the next Pen-F.
-
27 Aug 2026
Before pstack, Lauren Tan Wrote About Fear
Lauren Tan's 2017 essay on fear and learning: the failed startup, the self-taught start, the Netflix career, and the line I am still learning.
-
27 Aug 2026
Go Deep First: Notes on Lauren Tan's pstack
A reading of Lauren Tan's How I Use Cursor: agents as amnesiac new hires, verification as the bottleneck, pstack as playbooks, Benny as a work-in-progress factory. Editorial on what transfers off Cursor.
-
26 Aug 2026
Google's ATLAS: Mapping AI Use at the Scale of an Economy
15 million Gemini conversations mapped onto occupations, tasks, and household time — what the numbers support, which ones to discount, three interactive explainers, and an editorial on the new Solow paradox.
-
26 Aug 2026
The Family LLM Box: Mac mini vs. Mac Studio Over Two Years
Mac mini M6 32GB vs. M5 Pro 64GB vs. M5 Ultra 256GB: two-year cost of ownership against ChatGPT, Claude, Gemini, and Grok subscriptions for a household with kids.
-
26 Aug 2026
Portable Computer: Perplexity Puts the Whole Agent Runtime on Your Desk
Perplexity’s Portable Computer on NVIDIA DGX Spark: orchestrator, subagents, and agent harness run locally; escalation stays metered cloud. What changes when you own the runtime.
-
25 Aug 2026
If Every Agent Can Do the Work, Why Do You Still Need a System?
Roland Wayne’s thread read against Coase’s 1937 Economica essay: the cost of using the price mechanism, direction inside a firm, why firms do not grow without limit, and the API / CLI / access-token split.
-
25 Aug 2026
You Don’t Have to Say “Jobs” to Trigger Job Fear
A lab experiment: short videos about AI agency and control raised job-replacement anxiety even though none of them mentioned jobs.
-
24 Aug 2026
The Tiny Red Dot Lives in a Watershed We Haven’t Finished Mapping
The viral Laniakea map: a ~60% chance we sit in the larger Shapley basin of attraction — a drainage of galaxy flows, not a bound supercluster, and the outer rim is still off the page.
-
24 Aug 2026
Skills That Scan Clean Still Fail at Runtime: NVIDIA’s Skill Lift
ACES: structural skill scanners correlate with LLM-judge quality at ρ=0.14. Skill Lift — paired with/without-skill runs — is the metric that answers whether a skill actually helps.
-
24 Aug 2026
A Reading Path for RL on LLMs
Cameron Wolfe’s RL-for-LLMs bibliography as a path, plus a plain-English glossary of TRPO, PPO, GAE, KL, GRPO, and the rest of the alphabet soup.
-
24 Aug 2026
The AI Miracle Was a Straight Line
A LinkedIn-friendly tour of scaling laws: the log-log ruler, 6ND, Chinchilla, the emergence mirage, and why 2024 forked the line into train / think / efficiency.
-
24 Aug 2026
Outrunning Light: Cherenkov Radiation as an Electromagnetic Shock Front
From Mathelirium’s simulation: why a charged particle can outrun light inside a dielectric, how the Cherenkov cone forms, and why reactor pools glow blue.
-
22 Aug 2026
Prompt as Code in the GPT-Image-2 Prompt Library
A review of a structured prompt library, what its examples organize, and why structure can improve maintainability without making image generation deterministic.
-
22 Aug 2026
Graph Engineering: From 1 Prompt to 100 Agents Running in One System
An architectural masterclass by Hanako (@hanakoxbt): cutting imaginary edges, node contracts, the 4 canonical topologies, verifier gates, and durable state for massive multi-agent systems.
-
22 Aug 2026
Notes on Andrew Ng's AI Engineering Skills Map
Notes on a skills map derived from more than 10,000 job postings and dozens of interviews, with Leslie's interpretation kept separate from the source.
-
22 Aug 2026
A /eli5 Skill Pattern for Interactive HTML Explanations
A small Claude skill pattern for turning a topic into a self-contained HTML explanation, including an explicitly adapted SKILL.md example.
-
22 Aug 2026
FreeToken: Reported MoE Serving on Consumer Hardware
A review of FreeToken's reported consumer-hardware MoE serving approach, followed by a reproducible protocol for testing it locally.
-
21 Aug 2026
Why Some Early-Career Professionals Work Well with AI—and Others Don’t
Notes on a field study of 523 early-career professionals and its three observed patterns of AI use, including the limits of treating those clusters as fixed identities.
-
21 Aug 2026
Agency Agents: 140+ Open-Source Expert Personas
Notes on 140+ structured agent personas across 12 divisions, including how their role, scope, deliverable, and failure-mode prompts are organized.
-
21 Aug 2026
Codex as a Platform: Embedding OpenAI's Agent Harness
Notes on Codex exec, the SDK, app-server, and the approval boundaries needed when an agent is embedded in existing software.
-
21 Aug 2026
Transparent Backgrounds in the GPT-Image-2 API
The API parameters and base64 decoding needed for transparent PNG or WebP output, plus the cases that still need edge inspection.
-
21 Aug 2026
What Is an Agent Harness? One Practical Breakdown
One practical decomposition of an agent harness into instructions, tools, an execution loop, and the translation layer around model calls.
-
21 Aug 2026
Claude Academy: A Limited Note on the Launch
A deliberately limited inventory of what Anthropic's launch material supports, with direct links for checking the current curriculum.
-
20 Aug 2026
Autonomous Hexapod Architecture: Dual-Pi Compute, 40 TOPS Vision NPU, and Evaluating Modern Edge LLMs
Engineering blueprint for a Freenove Big Hexapod: Pi 5 + Pi Zero 2W partitioning, Hailo-10H 40 TOPS vision, 360° LIDAR SLAM, and upgrading from Gemma to Qwen 2.5 on-device.
-
20 Aug 2026
IP as Logo: A Minimalist Agent Skill for Generating High-Impact Mascot Icons
Highlighting s1dashu's open-source tool: enforcing strict geometric simplicity, 3-color palettes, and context-aware brand mascots inside AI coding agents.
-
20 Aug 2026
Moderna's Personalized mRNA Cancer Vaccine: Mechanism and Clinical Evidence
Technical breakdown of mRNA-4157/V940 neoantigen selection, PD-1 synergy, 5-year Phase 2 follow-up, and Phase 3 trial milestones.
-
20 Aug 2026
The River of Spacetime: Demystifying Black Holes & the Event Horizon
Interactive simulation and physical breakdown of Gullstrand–Painlevé spacetime waterfalls, light cone tipping, and event horizon mechanics.
-
20 Aug 2026
Ex-Vivo Neural Plasticity in a Closed-Loop Robotic Interface
A careful look at a reported closed-loop experiment connecting maintained human cortical tissue, a microelectrode array, and robotic piano-key actuation.
-
19 Aug 2026
A Local Model Candidate for a Single RTX 5090
A sourced checkpoint candidate for one RTX 5090, with caveats around reported throughput, context-memory fit, and reproducibility.
-
19 Aug 2026
Opening the log
Why this site exists: keeping a transparent, public record of works-in-progress rather than a finished portfolio.
Interactive Physics Laboratory
The River of Spacetime
A real-time relativistic simulation demystifying black holes and the event horizon based on Professor Brian Cox's spacetime waterfall and causal light-cone models.
Reading Pipeline & Mental Models
Currently on the Desk
Books actively shaping my technical architectures, historical perspectives, and philosophical foundations.
The Writers' Castle
Hands-On Large Language Models
Poems & Prayers
The Evolution of God
Active Initiatives
Research & Build Tracks
Four parallel experiments and long-term endeavors currently in progress.
Local Model Architectures
Investigating extreme low-bit quants (NVFP4, FP8, MTP speculative decoding) to achieve high-throughput (65–140+ tok/s) local agent inference on single GPUs.
Speculative Sci-Fi Novel
A long-form speculative fiction novel running parallel to technical systems. Deep world-building centered on distributed artificial intelligence and human agency.
Raspberry Pi Hexapod Agent ↗
An embodied AI robot on a Pi 5 + Pi Zero 2W. Bridging 40 TOPS Hailo-10H vision, 360° LIDAR SLAM, and on-device Qwen 2.5 for autonomous navigation.
Agentic Aircraft Leasing Platform ↗
A single-operator, agentic aircraft-leasing business model built to execute like a game loop. An experiment in high-autonomy multi-agent orchestration — playable live as Lobster AeroCorp OS.
Setup & Tooling
Lab Environment
Hardware rigs and software runtimes used across active research tracks.
Primary GPU Rig
- Current GPURTX 4090 (24GB Ada)
- Target ArchitectureRTX 5090 (Blackwell NVFP4)
- Host OSLinux / Ubuntu 24.04 LTS
- CUDA / DriverCUDA 12.8+ / 570.xx
Inference & Quantization
- EnginesvLLM, llama.cpp, Ollama
- Quant ToolNVIDIA ModelOpt
- FormatsGGUF, NVFP4, FP8, AWQ
- AccelerationFlashAttention-3, MTP
Embedded & Stack
- Edge ComputeRaspberry Pi 5 (8GB)
- LanguagesPython, TypeScript, C++
- Agentic FrameworkCustom Lightweight Loops
- Web InfrastructureNode.js Express / Static Edge
About
Leslie Li
I was a software engineer years ago. These days I am a hobbyist — tech, astronomy, quantum physics, philosophy, and psychology. I also shoot street photography.
leslieli.dev is a public log of what I am researching, building, and writing.
If you want to reach me, use LinkedIn.
Far away from the black hole, the river of space flows gently inwards at subluminal speeds. Light and rockets can easily paddle upstream and escape into deep space.