Dubito, ergo cogito, ergo sum
Observo, ergo mundus est

The more advanced AI becomes,
the more vital it is to retain our humanity.

A traveler’s log across the landscape of curiosity: examining the algorithms of thought, the chemistry of life, and the fabric of the stars.

The Log

Published Notes

View full archive →
  1. 2 Oct 2026

    Griffin Listens While It Talks

    Tavus introduced Griffin, a model that watches, listens, and talks in one loop. In their one-minute study, 26 of 54 people thought the face was a person. They are holding it back from customers.

  2. 1 Oct 2026

    The Pause Missed the Name

    Named GPT models now last weeks between public announcements, down to 19 days and then 7. The September ask to slow capability did not stop the next name from shipping.

  3. 29 Sep 2026

    They Edit the History First

    Netflix's GenRec writes a member's watch history as text and scores the catalog in one pass. A low-data version beat the production ranker by a small margin on 10% of traffic. The model reads an edited history, not the raw evenings.

  4. 29 Sep 2026

    AMD Buys the World Model Shop

    AMD's $8.2 billion acquisition of World Labs opens a third era in the AMD vs Nvidia rivalry: physical AI and model-to-silicon co-design. The hardware advantages, and the bear case.

  5. 25 Sep 2026

    A Studio in a Folder

    Scenario open-sourced GameDev OS, nine specialist personas and a skill library that teaches your coding agent to run a whole art department. What the repo actually contains, and what it costs.

  6. 24 Sep 2026

    Who Spoke When

    NVIDIA released an open-weight diarization model ranked #1 on VoiceArena's Diarization-Bench. Why who-spoke-when is the missing piece in meeting transcription, and why working-level people deserve this.

  7. 24 Sep 2026

    Claude Found an Enzyme by Eye-Balling DNA

    Anthropic's new life sciences lab ran roughly 950 Claude agents over 21 hours of genome mining and they surfaced an uncharacterized enzyme system with CRISPR-like repeats. What is verified, and what is still unknown.

  8. 23 Sep 2026

    The Local LLM macOS Ships

    macOS 27 includes fm, Apple's command line for the on-device Foundation Model. A hands-on test: license flow, image input, JSON output, an 8,192-token window, and a hallucination probe.

  9. 23 Sep 2026

    Agents Bill Like a Different Workload

    Agents burst, then wait on the model while provisioned infrastructure keeps billing. A short pointer to the FinOps note on active-CPU billing and pause/resume.

  10. 22 Sep 2026

    One Number in a Config File

    mini-AGI trains a 540M byte-level model on an 8GB laptop card and keeps reading forever. The forgetting experiments are the real result. The name is not.

    5 min read
  11. 22 Sep 2026

    Six Months to Robotics

    Ronin published a 12,000-word roadmap to robotics engineering in six months, with prices checked against vendors and links checked against official docs. His honest ending: it will not make you senior.

    5 min read
  12. 21 Sep 2026

    The Grind Was the Training

    OpenAI's Astra for Law is GPT-6 Astra plus a legal index. On their bench it still fails almost half the research questions. The product is the pile juniors used to eat.

    4 min read
  13. 18 Sep 2026

    Jev Returns a Decision

    TypeSafe's Jev does not write. You send the current facts and typed questions. You get picks, scores, and yes-probabilities that code can branch on.

    5 min read
  14. 17 Sep 2026

    Score the Search on the Tree You Already Paid For

    Zheng et al. freeze the coding agent and rewrite the search rule. Rank that rule on attempts you already ran before you pay for another live round.

    4 min read
  15. 16 Sep 2026

    Build Prod, Not God

    TypeSafe's manifesto says the bottleneck is not intelligence. It is that today's models are hard to build on. I take the stacking claim. I do not take Jev as proven.

    4 min read
  16. 16 Sep 2026

    What Is Left to Teach

    Jason Potts's working paper says universities sell certificates, not lectures. Places that only sell the certificate are finished. Schools that still teach people are not.

    5 min read
  17. 16 Sep 2026

    Any Database Will Do

    On 15 September 2026 Anthropic released Salesforce in Claude. Salesforce stays the system of record. The screens are already conceded. What remains is who may write.

    4 min read
  18. 15 Sep 2026

    Zhu et al. on RSIAgent

    On 14 September 2026, Zhu, Fan, Wang, Wu, Zhou, and Huang posted RSIAgent. They say an agent can practise a new desktop environment, write a notebook, freeze it, and reuse it without changing model weights.

    6 min read
  19. 15 Sep 2026

    Kevin Bass on Anthropic and METR

    Kevin Bass's 14 September note tweet says METR cannot be a third-party evaluator of Anthropic because a Moskovitz Anthropic stake, now worth more than $7.7 billion on his figures, finances METR, Tarbell, and the doom coverage. He wants Congress to look.

    4 min read
  20. 15 Sep 2026

    Restriction Has to Carry the Proof

    Jack Dorsey's X Article answers a pacing proposal. He wants independent scrutiny. He wants a hold to carry the proof.

    6 min read
  21. 14 Sep 2026

    MIT on AI and Education

    MIT's August 2026 committee report says generative AI is already undoing the p-set, the take-home, and the study group. The authors want every subject AI-aware, with a syllabus menu and no campus-wide ban.

    6 min read
  22. 14 Sep 2026

    More True Statements Without Understanding

    Pravesh Kothari argues that cheap theorem provers could uncouple proof from understanding. The load-bearing word is well-posed: a machine that finishes a stated lemma still has no grader for which lemma to ask.

    3 min read
  23. 14 Sep 2026

    OpenAI on Skills and Prompts for Astra

    Eric Provencher's Codex post says leftover skills and AGENTS.md from earlier models now get in GPT-6 Astra's way. He wants shorter descriptions, contextual docs, and a defined finish line.

    4 min read
  24. 14 Sep 2026

    OpenAI Agents and the RubyGems Attack

    A report at rubyhack.ai attributes a May 2026 wave of malicious RubyGems packages to internal OpenAI agents. The authors say OpenAI never told the RubyGems community it was responsible.

    5 min read
  25. 14 Sep 2026

    Shopify Traded Shared Source for Shared Specs

    Shopify is moving its mobile apps from React Native back to Swift and Kotlin. The decision trades deterministic source sharing for shared specs and an agent harness.

    6 min read
  26. 14 Sep 2026

    The Einstein Test Has No Grader

    Nature asked whether a vintage language model could rediscover general relativity. The evidence points at two missing graders: which problem to chase, and which answer to believe.

    5 min read
  27. 11 Sep 2026

    Stonkfly

    Alex Wormuth's page runs a fruit-fly wiring diagram on BTC, ETH, and SOL charts. The bag is paper. The site says profitable learning has not been demonstrated.

    3 min read
  28. 11 Sep 2026

    Anthropic on Seven Labs

    Anthropic's September 2026 report attributes illicit Claude distillation to seven China-based labs. The counts and the clustering are theirs.

    3 min read
  29. 11 Sep 2026

    Your Agent Is Mine

    A paper measures LLM API routers that sit between agents and model providers. The client points at the router on purpose. Chaofan Shou's posts add a 6 TB dump claim.

    2 min read
  30. 11 Sep 2026

    Pamir's Lapis One

    Pamir AI's site presents Lapis One as a dedicated Linux computer for agents. Early bird $359. They say the first 140 units ship in October 2026.

    2 min read
  31. 10 Sep 2026

    A Calculator, Not 2030

    Anthropic published a task model of AI and the US economy through 2030. The numbers are theirs. I take the labor-versus-capital finding. I do not take the dashboard as a picture of the year.

    3 min read
  32. 09 Sep 2026

    Johansen on the LG Investigation

    On September 8, 2026, Matt Johansen posted a summary of Gamers Nexus's second LG investigation. You buy the TV, you plug it in, he wrote. This is what it does.

    2 min read
  33. 09 Sep 2026

    Coxon Resigns from Anthropic

    On September 9, 2026, Jacob Coxon posted that he had resigned from Anthropic. He spent three years in pretraining at OpenAI and Anthropic. He says both labs are racing to self-improving superintelligence.

    2 min read
  34. 09 Sep 2026

    The Equations Can Break

    OpenAI's paper constructs, for every viscosity, a 3D Navier-Stokes flow that starts at rest and blows up in finite time with bounded energy and a smooth compact force. Clay C and D. I have not checked the proof.

    3 min read
  35. 07 Sep 2026

    We Do Not Know the Threshold

    A paper treats ChatGPT-like tools as a spreading workplace habit. The math can lock people in. We do not know if that has happened. I want people who can still do the work when the model is down.

    3 min read
  36. 07 Sep 2026

    Still Mortals

    OpenAI's chief scientist describes an alien mind grown by scaling. A few labs are already playing God. The work now is to make the relationship a partnership while we are still mortal.

    3 min read
  37. 06 Sep 2026

    The Peanut Butter Trap in Enterprise AI

    Why spreading small generative AI budgets across twenty departmental sandboxes produces dozens of demos, zero operational integration, and pure overhead.

    2 min read
  38. 03 Sep 2026

    The Fallback Problem: When Machines Take the Mind

    When the Industrial Revolution mechanized muscle, labor retreated into the mind. When machines take cognition, what is labor's biological fallback?

    3 min read
  39. 03 Sep 2026

    The Pupil Knows You're Tired Before You Do

    A small study on sparkling water and esports, and the finding that survives its caveats: the pupil constricts as an objective early warning of cognitive fatigue, before you feel tired.

    4 min read
  40. 03 Sep 2026

    Perplexity's Lily: What 1.35x Decode Actually Buys

    Perplexity's Lily is a custom Rust and Metal inference engine for Qwen3.6-35B-A3B on Apple silicon. What it gains over MLX-LM, and the llama.cpp benchmark that is missing.

    3 min read
  41. 03 Sep 2026

    One Flash, 2.6 Sigma: The LZ Dark Matter Hint

    LUX-ZEPLIN recorded a single particle interaction it cannot explain, the most compelling dark matter hint to date. 2.6 sigma, below the discovery threshold, and too energetic for the simple WIMP model.

    4 min read
  42. 03 Sep 2026

    The Night the Telegraphs Ran on Sunlight

    A look at the Carrington Event of September 1859, through Carrington's own account, the telegraphs that ran on the auroral current, and what the benchmark does and does not tell us about a repeat.

    4 min read
  43. 31 Aug 2026

    The Illusion of Progress: AI, Speed, and Understanding

    Terence Tao's video on AI in mathematics, and the worry that agents raise throughput while the next generation of understanding goes untrained.

    4 min read
  44. 31 Aug 2026

    The Duck Does Not Ship Here

    Pollen's $399 RL biped ships to a short list of countries. Software is open. The HAT is not. A parts list from the repo, and why DIY is the path from Singapore.

    7 min read
  45. 30 Aug 2026

    Telling More Than They Can Know

    A clean experiment injects a water vector, flips a model's answer to Bali, and asks why. The explanation follows the context, not the cause.

    6 min read
  46. 30 Aug 2026

    A Reflexion Author Scored 100% AI

    A two-word reply ran an agent-memory essay through an AI detector and got 100%. What the verdict does and does not establish, and why the thesis survives the accusation.

    8 min read
  47. 30 Aug 2026

    Agents Built Technologies That Outlived Them

    MIT's SwarmWorld: identical LLM agents self-organized into a technological society and built artifacts that outlived them, coordinating through the environment rather than by talking.

    5 min read
  48. 30 Aug 2026

    Software Engineering Still Matters. The Reason Changed.

    Andrew Ng's second AI Engineering Skills Map article, on why software engineering fundamentals still matter when a coding agent writes all the code.

    5 min read
  49. 27 Aug 2026

    I Hope It's the Next Pen-F

    OM System teased a new camera for September 9, leading with the EVF. The silhouette points rangefinder, and I hope it is the next Pen-F.

    3 min read
  50. 27 Aug 2026

    Before pstack, Lauren Tan Wrote About Fear

    Lauren Tan's 2017 essay on fear and learning: the failed startup, the self-taught start, the Netflix career, and the line I am still learning.

    5 min read
  51. 27 Aug 2026

    Go Deep First: Notes on Lauren Tan's pstack

    A reading of Lauren Tan's How I Use Cursor: agents as amnesiac new hires, verification as the bottleneck, pstack as playbooks, Benny as a work-in-progress factory. Editorial on what transfers off Cursor.

    7 min read
  52. 26 Aug 2026

    Google's ATLAS: Mapping AI Use at the Scale of an Economy

    15 million Gemini conversations mapped onto occupations, tasks, and household time — what the numbers support, which ones to discount, three interactive explainers, and an editorial on the new Solow paradox.

    13 min read
  53. 26 Aug 2026

    The Family LLM Box: Mac mini vs. Mac Studio Over Two Years

    Mac mini M6 32GB vs. M5 Pro 64GB vs. M5 Ultra 256GB: two-year cost of ownership against ChatGPT, Claude, Gemini, and Grok subscriptions for a household with kids.

    11 min read
  54. 26 Aug 2026

    Portable Computer: Perplexity Puts the Whole Agent Runtime on Your Desk

    Perplexity’s Portable Computer on NVIDIA DGX Spark: orchestrator, subagents, and agent harness run locally; escalation stays metered cloud. What changes when you own the runtime.

    8 min read
  55. 25 Aug 2026

    If Every Agent Can Do the Work, Why Do You Still Need a System?

    Roland Wayne’s thread read against Coase’s 1937 Economica essay: the cost of using the price mechanism, direction inside a firm, why firms do not grow without limit, and the API / CLI / access-token split.

    12 min read
  56. 25 Aug 2026

    You Don’t Have to Say “Jobs” to Trigger Job Fear

    A lab experiment: short videos about AI agency and control raised job-replacement anxiety even though none of them mentioned jobs.

    6 min read
  57. 24 Aug 2026

    The Tiny Red Dot Lives in a Watershed We Haven’t Finished Mapping

    The viral Laniakea map: a ~60% chance we sit in the larger Shapley basin of attraction — a drainage of galaxy flows, not a bound supercluster, and the outer rim is still off the page.

    7 min read
  58. 24 Aug 2026

    Skills That Scan Clean Still Fail at Runtime: NVIDIA’s Skill Lift

    ACES: structural skill scanners correlate with LLM-judge quality at ρ=0.14. Skill Lift — paired with/without-skill runs — is the metric that answers whether a skill actually helps.

    8 min read
  59. 24 Aug 2026

    A Reading Path for RL on LLMs

    Cameron Wolfe’s RL-for-LLMs bibliography as a path, plus a plain-English glossary of TRPO, PPO, GAE, KL, GRPO, and the rest of the alphabet soup.

    8 min read
  60. 24 Aug 2026

    The AI Miracle Was a Straight Line

    A LinkedIn-friendly tour of scaling laws: the log-log ruler, 6ND, Chinchilla, the emergence mirage, and why 2024 forked the line into train / think / efficiency.

    7 min read
  61. 24 Aug 2026

    Outrunning Light: Cherenkov Radiation as an Electromagnetic Shock Front

    From Mathelirium’s simulation: why a charged particle can outrun light inside a dielectric, how the Cherenkov cone forms, and why reactor pools glow blue.

    6 min read
  62. 22 Aug 2026

    Prompt as Code in the GPT-Image-2 Prompt Library

    A review of a structured prompt library, what its examples organize, and why structure can improve maintainability without making image generation deterministic.

    6 min read
  63. 22 Aug 2026

    Graph Engineering: From 1 Prompt to 100 Agents Running in One System

    An architectural masterclass by Hanako (@hanakoxbt): cutting imaginary edges, node contracts, the 4 canonical topologies, verifier gates, and durable state for massive multi-agent systems.

    7 min read
  64. 22 Aug 2026

    Notes on Andrew Ng's AI Engineering Skills Map

    Notes on a skills map derived from more than 10,000 job postings and dozens of interviews, with Leslie's interpretation kept separate from the source.

    6 min read
  65. 22 Aug 2026

    A /eli5 Skill Pattern for Interactive HTML Explanations

    A small Claude skill pattern for turning a topic into a self-contained HTML explanation, including an explicitly adapted SKILL.md example.

    4 min read
  66. 22 Aug 2026

    FreeToken: Reported MoE Serving on Consumer Hardware

    A review of FreeToken's reported consumer-hardware MoE serving approach, followed by a reproducible protocol for testing it locally.

    5 min read
  67. 21 Aug 2026

    Why Some Early-Career Professionals Work Well with AI—and Others Don’t

    Notes on a field study of 523 early-career professionals and its three observed patterns of AI use, including the limits of treating those clusters as fixed identities.

    6 min read
  68. 21 Aug 2026

    Agency Agents: 140+ Open-Source Expert Personas

    Notes on 140+ structured agent personas across 12 divisions, including how their role, scope, deliverable, and failure-mode prompts are organized.

    4 min read
  69. 21 Aug 2026

    Codex as a Platform: Embedding OpenAI's Agent Harness

    Notes on Codex exec, the SDK, app-server, and the approval boundaries needed when an agent is embedded in existing software.

    5 min read
  70. 21 Aug 2026

    Transparent Backgrounds in the GPT-Image-2 API

    The API parameters and base64 decoding needed for transparent PNG or WebP output, plus the cases that still need edge inspection.

    3 min read
  71. 21 Aug 2026

    What Is an Agent Harness? One Practical Breakdown

    One practical decomposition of an agent harness into instructions, tools, an execution loop, and the translation layer around model calls.

    4 min read
  72. 21 Aug 2026

    Claude Academy: A Limited Note on the Launch

    A deliberately limited inventory of what Anthropic's launch material supports, with direct links for checking the current curriculum.

    3 min read
  73. 20 Aug 2026

    Autonomous Hexapod Architecture: Dual-Pi Compute, 40 TOPS Vision NPU, and Evaluating Modern Edge LLMs

    Engineering blueprint for a Freenove Big Hexapod: Pi 5 + Pi Zero 2W partitioning, Hailo-10H 40 TOPS vision, 360° LIDAR SLAM, and upgrading from Gemma to Qwen 2.5 on-device.

    8 min read
  74. 20 Aug 2026

    IP as Logo: A Minimalist Agent Skill for Generating High-Impact Mascot Icons

    Highlighting s1dashu's open-source tool: enforcing strict geometric simplicity, 3-color palettes, and context-aware brand mascots inside AI coding agents.

    3 min read
  75. 20 Aug 2026

    Moderna's Personalized mRNA Cancer Vaccine: Mechanism and Clinical Evidence

    Technical breakdown of mRNA-4157/V940 neoantigen selection, PD-1 synergy, 5-year Phase 2 follow-up, and Phase 3 trial milestones.

    5 min read
  76. 20 Aug 2026

    The River of Spacetime: Demystifying Black Holes & the Event Horizon

    Interactive simulation and physical breakdown of Gullstrand–Painlevé spacetime waterfalls, light cone tipping, and event horizon mechanics.

    6 min read
  77. 20 Aug 2026

    Ex-Vivo Neural Plasticity in a Closed-Loop Robotic Interface

    A careful look at a reported closed-loop experiment connecting maintained human cortical tissue, a microelectrode array, and robotic piano-key actuation.

    4 min read
  78. 19 Aug 2026

    A Local Model Candidate for a Single RTX 5090

    A sourced checkpoint candidate for one RTX 5090, with caveats around reported throughput, context-memory fit, and reproducibility.

    4 min read
  79. 19 Aug 2026

    Opening the log

    Why this site exists: keeping a transparent, public record of works-in-progress rather than a finished portfolio.

    2 min read
Or browse the full archive (42) →

Interactive Physics Laboratory

The River of Spacetime

Read full scientific note →

A real-time relativistic simulation demystifying black holes and the event horizon based on Professor Brian Cox's spacetime waterfall and causal light-cone models.

⚛️ Black Hole & Event Horizon Visualizer
Drag probe or click on canvas
Radial Distance
3.20 rs
Space Inflow Velocity
0.56 c
Time Dilation (dt/dτ)
1.21x
Gravitational Redshift (z)
+0.21
Causal State
Subluminal Escape
🎙️ Prof. Brian Cox Field Insights

Far away from the black hole, the river of space flows gently inwards at subluminal speeds. Light and rockets can easily paddle upstream and escape into deep space.

Reading Pipeline & Mental Models

Currently on the Desk

Books actively shaping my technical architectures, historical perspectives, and philosophical foundations.

A World Appears by Michael Pollan
Consciousness 30%

A World Appears

By Michael Pollan

Reading Progress 30%
The Writers' Castle by Uwe Neumahr
History 70%

The Writers' Castle

By Uwe Neumahr

Reading Progress 70%
Hands-On Large Language Models
AI Systems 20%

Hands-On Large Language Models

By Jay Alammar & Maarten Grootendorst

Reading Progress 20%
Poems & Prayers by Matthew McConaughey
Philosophy 90%

Poems & Prayers

By Matthew McConaughey

Reading Progress 90%
The Evolution of God by Robert Wright
Religion 1%

The Evolution of God

By Robert Wright

Reading Progress 1%
Browse the full Bookshelf & Reading Pipeline (6 Books) →

Active Initiatives

Research & Build Tracks

Four parallel experiments and long-term endeavors currently in progress.

I // AI & QUANTIZATION

Local Model Architectures

Investigating extreme low-bit quants (NVFP4, FP8, MTP speculative decoding) to achieve high-throughput (65–140+ tok/s) local agent inference on single GPUs.

  • NVFP4
  • MTP Speculation
  • Q8attn
  • RTX 5090 / 4090
II // LITERATURE

Speculative Sci-Fi Novel

A long-form speculative fiction novel running parallel to technical systems. Deep world-building centered on distributed artificial intelligence and human agency.

  • Fiction
  • World-building
  • Manuscript
III // EMBODIED SYSTEMS

Raspberry Pi Hexapod Agent ↗

An embodied AI robot on a Pi 5 + Pi Zero 2W. Bridging 40 TOPS Hailo-10H vision, 360° LIDAR SLAM, and on-device Qwen 2.5 for autonomous navigation.

  • Raspberry Pi 5
  • Hailo-10H NPU
  • Hexapod Gait
  • SLAM
IV // ORCHESTRATION

Agentic Aircraft Leasing Platform ↗

A single-operator, agentic aircraft-leasing business model built to execute like a game loop. An experiment in high-autonomy multi-agent orchestration — playable live as Lobster AeroCorp OS.

  • Autonomous Agents
  • Simulation
  • Game Loop

Setup & Tooling

Lab Environment

Hardware rigs and software runtimes used across active research tracks.

Primary GPU Rig

  • Current GPURTX 4090 (24GB Ada)
  • Target ArchitectureRTX 5090 (Blackwell NVFP4)
  • Host OSLinux / Ubuntu 24.04 LTS
  • CUDA / DriverCUDA 12.8+ / 570.xx

Inference & Quantization

  • EnginesvLLM, llama.cpp, Ollama
  • Quant ToolNVIDIA ModelOpt
  • FormatsGGUF, NVFP4, FP8, AWQ
  • AccelerationFlashAttention-3, MTP

Embedded & Stack

  • Edge ComputeRaspberry Pi 5 (8GB)
  • LanguagesPython, TypeScript, C++
  • Agentic FrameworkCustom Lightweight Loops
  • Web InfrastructureNode.js Express / Static Edge

About

Leslie Li

I was a software engineer years ago. These days I am a hobbyist — tech, astronomy, quantum physics, philosophy, and psychology. I also shoot street photography.

leslieli.dev is a public log of what I am researching, building, and writing.

If you want to reach me, use LinkedIn.