<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>dotmd</title><description>AI agents and startup lessons — by Timi.</description><link>https://blog.timi.click/</link><language>en-us</language><item><title>Shipping on the Go</title><link>https://blog.timi.click/shipping-on-the-go/</link><guid isPermaLink="true">https://blog.timi.click/shipping-on-the-go/</guid><description>A step-by-step guide to running your coding agents on an always-on VPS and driving them from your laptop and phone with T3 Code.</description><pubDate>Sat, 10 Oct 2026 00:00:00 GMT</pubDate><content:encoded>import Callout from &apos;../../components/guide/Callout.astro&apos;;
import FleetMap from &apos;../../components/guide/FleetMap.astro&apos;;
import ProxyFailover from &apos;../../components/guide/ProxyFailover.astro&apos;;
import ConnectRoutes from &apos;../../components/guide/ConnectRoutes.astro&apos;;

&lt;Callout type=&quot;tip&quot; title=&quot;TL;DR&quot;&gt;
You don&apos;t have to read all of this. Send this link to your agent and ask it to walk you through the setup: [blog.timi.click/shipping-on-the-go.llm.txt](https://blog.timi.click/shipping-on-the-go.llm.txt)
&lt;/Callout&gt;

People keep asking me what my setup looks like, so this is the full guide.

My agents don&apos;t run on my laptop anymore. They run on a cheap VPS that never sleeps, and I control them from my laptop and my phone. I can start a task at my desk, close the lid, and check on it from my phone while I&apos;m out.

The biggest win for me is coding from my phone on the train. I don&apos;t have to pull my laptop out and look like a creep.

&lt;FleetMap /&gt;

The setup has seven steps:

1. **[T3 Code](https://t3.codes) on your laptop**, the app you&apos;ll drive everything from.
2. **A VPS**, the always-on machine your agents live on.
3. **A VPS environment that matches your laptop**, so agents can do the same work there.
4. **CLI Proxy** *(optional)*, if you have more than one subscription on a provider.
5. **T3 Code on the VPS**, connected to your laptop with full permissions.
6. **Your laptop on the proxy** *(optional)*, so it shares the same pool of accounts.
7. **T3 Code on your phone**, connected to both machines.

Once your agent has SSH access to the VPS, it can do most of this for you. I point out where that helps.

## 1. Install T3 Code on your laptop

[T3 Code](https://github.com/pingdotgg/t3code) is a control surface for coding agents. It doesn&apos;t replace Claude Code or Codex. It runs them for you, using the subscriptions you&apos;re already logged into, and gives you one app to manage threads, worktrees, and machines.

Download the desktop app from the [GitHub releases page](https://github.com/pingdotgg/t3code/releases). I recommend the **nightly** build (the releases tagged `-nightly`). It moves fast, and the newest features, including a lot of the remote-machine work in this guide, land there first. On macOS you can also install it with Homebrew: `brew install --cask t3-code@nightly`.

Before you open it, install at least one agent CLI on your laptop and log in:

```bash
# Claude Code
curl -fsSL https://claude.ai/install.sh | bash
claude auth login

# Codex
npm install -g @openai/codex
codex login
```

Open T3 Code, start a thread in any project, and check that your agent answers.

## 2. Get a VPS

Any Linux VPS works. I use [Contabo](https://contabo.com) because it&apos;s cheap and reliable. [Hetzner](https://www.hetzner.com/cloud) is a good alternative, and any other provider works too.

What I&apos;d look for:

- **Ubuntu 24.04**, so every command in this guide works as written.
- **At least 8 GB of RAM.** Agents, language servers, Docker, and a headless browser add up quickly.
- **A region close to you**, so the terminal feels responsive.

Once it&apos;s running, put it on [Tailscale](https://tailscale.com). Tailscale gives your laptop, phone, and VPS a private network. Every device gets a stable address and a name, and nothing has to be exposed to the public internet.

```bash
# on the VPS
curl -fsSL https://tailscale.com/install.sh | sh
sudo tailscale up
```

Install Tailscale on your laptop and phone too, and sign in with the same account. From now on, you can reach the VPS by its tailnet name from anywhere.

&lt;Callout type=&quot;tip&quot;&gt;
Once you can SSH into the VPS from your laptop, you can hand most of the next step to your local agent. Tell it what you use locally, and ask it to install the same tools on the server.
&lt;/Callout&gt;

## 3. Make the VPS feel like your laptop

Your agents on the VPS can only do what the VPS can do. If your laptop has Node, Bun, Python, Docker, and the GitHub CLI, the VPS needs them too. Otherwise the first `npm install` or `docker compose up` an agent runs will fail.

Go through what you actually use. For me that&apos;s roughly:

```bash
# basics
sudo apt update &amp;&amp; sudo apt install -y git build-essential curl unzip

# GitHub CLI, then log in so agents can push and open PRs
sudo apt install -y gh
gh auth login

# Node (via fnm), Bun, and uv for Python
curl -fsSL https://fnm.vercel.app/install | bash
curl -fsSL https://bun.sh/install | bash
curl -LsSf https://astral.sh/uv/install.sh | sh

# Docker
curl -fsSL https://get.docker.com | sh

# the agent CLIs, logged in just like on your laptop
curl -fsSL https://claude.ai/install.sh | bash
npm install -g @openai/codex
```

Then log the agent CLIs in (`claude auth login`, `codex login`), clone the repos you work on, and copy over any `.env` files your projects need.

&lt;Callout type=&quot;gotcha&quot; title=&quot;Gotcha: old Codex builds&quot;&gt;
Install Codex from npm, not snap. The snap package was an older build for me, and it silently ignored custom providers, which breaks the proxy setup in the next step.
&lt;/Callout&gt;

&lt;Callout type=&quot;gotcha&quot; title=&quot;Gotcha: running as root&quot;&gt;
Most VPSes give you a root user. Claude Code refuses to skip permission prompts as root unless it knows it&apos;s on a sandboxed machine. If you want T3 to run Claude in full-access mode on the server, set `IS_SANDBOX=1` for the T3 service (I&apos;ll show where in step 5).
&lt;/Callout&gt;

## 4. Set up CLI Proxy (optional)

&lt;Callout type=&quot;optional&quot;&gt;
Skip this if you have one subscription per provider. Come back if you ever add a second.
&lt;/Callout&gt;

I have two Claude subscriptions and a Codex one. Before the proxy, I logged each machine into one account. When that account hit its five-hour limit, I logged out and logged into the other one by hand.

[CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) fixes that. You log every account into it once, and it gives you a single endpoint that speaks the Anthropic and OpenAI APIs. Your tools point at that endpoint instead of at Anthropic or OpenAI directly. When one account hits its limit, the proxy cools it down and retries on the next one.

&lt;ProxyFailover /&gt;

I run it in Docker on the VPS. Publish the port only on the VPS&apos;s **Tailscale address**. Then only your own devices can reach it.

```yaml
# /opt/cliproxy/docker-compose.yml
services:
  cli-proxy-api:
    image: eceasy/cli-proxy-api:latest
    ports: [&quot;&lt;vps-tailscale-ip&gt;:8317:8317&quot;]
    volumes:
      - ./config.yaml:/CLIProxyAPI/config.yaml
      - ./auths:/root/.cli-proxy-api
    restart: unless-stopped
```

In `config.yaml`, set an API key for your own tools under `access.api-keys`, and a `management.secret-key` for the dashboard. Then log each account in, following the [CLIProxyAPI README](https://github.com/router-for-me/CLIProxyAPI). The dashboard lives at `http://&lt;vps-tailscale-name&gt;:8317/management.html` and shows every account, its requests, and its remaining quota.

Now point the VPS&apos;s agents at it. Claude Code reads its environment from `~/.claude/settings.json`:

```json
{
  &quot;env&quot;: {
    &quot;ANTHROPIC_BASE_URL&quot;: &quot;http://&lt;vps-tailscale-name&gt;:8317&quot;,
    &quot;ANTHROPIC_API_KEY&quot;: &quot;&lt;your proxy api key&gt;&quot;
  }
}
```

Codex gets a custom provider in `~/.codex/config.toml`:

```toml
model_provider = &quot;fleet&quot;

[model_providers.fleet]
name = &quot;Fleet proxy&quot;
base_url = &quot;http://&lt;vps-tailscale-name&gt;:8317/v1&quot;
env_key = &quot;FLEET_PROXY_KEY&quot;  # an env var holding your proxy api key
wire_api = &quot;responses&quot;
```

Run a prompt with each CLI and check the proxy dashboard. If the request count goes up, it&apos;s working.

&lt;Callout type=&quot;tip&quot; title=&quot;Not into proxies?&quot;&gt;
T3 Code can also run several Claude or Codex accounts natively, as separate provider instances with their own config directories. It doesn&apos;t fail over automatically, but it&apos;s less to run. See the [multiple accounts docs](https://github.com/pingdotgg/t3code/blob/main/docs/user/providers-claude.md).
&lt;/Callout&gt;

## 5. Run T3 Code on the VPS and connect to it

Install the T3 server on the VPS. The install script takes a channel, so you can match the nightly desktop app:

```bash
curl -fsSL https://t3.codes/install.sh | T3CODE_CHANNEL=nightly sh

# keep it running in the background, across reboots
t3 service install
```

Because T3 runs the CLIs you installed in step 3, it uses the proxy from step 4 automatically. You don&apos;t need to configure T3 for it.

Next, connect your laptop to it. There are two routes, and you can use both:

&lt;ConnectRoutes /&gt;

**Route A: Tailscale.** On the VPS, run:

```bash
t3 pair --tailscale
```

This publishes the server over HTTPS on your tailnet and prints a one-time pairing link and QR code. In the desktop app, go to **Settings → Connections → Add environment** and paste the link.

**Route B: T3 Connect.** On the VPS, run `t3 connect` and follow the sign-in. On your laptop, sign in to the same T3 Connect account in **Settings → Connections** and pick the environment. No Tailscale needed.

&lt;Callout type=&quot;gotcha&quot; title=&quot;Gotcha: limited permissions&quot;&gt;
Sometimes the app says your connection has limited permissions, and the file and browser panels don&apos;t work. This means the pairing gave your device only some scopes. Check under **Settings → Connections → Permissions**. To fix it, create a new link with `t3 pair --tailscale` and pair again. Its default scopes include files, settings, providers, and the browser. Then remove the old entry.
&lt;/Callout&gt;

Two more things I had to fix on the server before everything worked:

**The browser.** T3 can give agents a headless browser on the VPS, but it needs Chrome and its libraries installed first:

```bash
t3 browser setup
```

On a root VPS, Chrome&apos;s sandbox still won&apos;t start, so disable it for the T3 service. The same override is where `IS_SANDBOX=1` from step 3 goes:

```bash
systemctl --user edit t3code
```

```ini
[Service]
Environment=T3CODE_SERVER_BROWSER_SANDBOX=0
Environment=IS_SANDBOX=1
```

```bash
t3 service restart
```

Now start a thread on the VPS environment from your laptop. Ask the agent to run `hostname`. If it prints the VPS&apos;s name, you&apos;re driving the server.

## 6. Put your laptop on the proxy too (optional)

If you set up the proxy, point your laptop&apos;s CLIs at it as well, so local threads draw from the same pool of accounts. Use the same `~/.claude/settings.json` and `~/.codex/config.toml` from step 4 on your laptop, and set `FLEET_PROXY_KEY` in your environment.

Restart T3 Code afterwards so the agents pick up the new settings. The proxy dashboard should show requests from both machines.

&lt;Callout type=&quot;gotcha&quot; title=&quot;Gotcha: the proxy needs Tailscale&quot;&gt;
The proxy only listens on your tailnet. If Tailscale is off on your laptop, local agents can&apos;t reach it and will fail. Leave Tailscale running, or remove the proxy settings when you go offline.
&lt;/Callout&gt;

## 7. Put T3 Code on your phone

If you followed my advice and installed the nightly, you need the beta phone app. The App Store and Play Store versions can&apos;t connect to nightly builds.

- **iPhone and iPad:** join the [TestFlight beta](https://testflight.apple.com/join/XgaxaRtd).
- **Android:** join the [beta group](https://groups.google.com/g/t3-code-v2-beta), then open the [Play testing page](https://play.google.com/apps/testing/com.t3tools.t3code) with the same Google account and become a tester.

The nightly desktop app also shows these links as QR codes under **Settings → General → Mobile app**. On the stable build, use the regular [iOS](https://apps.apple.com/us/app/t3-code-remote-claude-more/id6787819824) or [Android](https://play.google.com/store/apps/details?id=com.t3tools.t3code) app.

Connect it the same way as your laptop. With Tailscale, run `t3 pair --tailscale` on the VPS and scan the QR code from **Settings → Environments** in the app. With T3 Connect, sign in. To reach your laptop&apos;s local threads from your phone, create a link in **Settings → Connections** on the desktop app.

Now you can start a long task, leave your desk, and check it from your phone. You can approve actions, send follow-ups, and start new threads from the app.

## The result

- Your agents run on a machine that never sleeps.
- Your laptop and phone can both start, follow, and steer them.
- With the proxy, every tool on every machine draws from every subscription you pay for. When one account hits its limit, the next one takes over.

If the VPS breaks, you can rebuild it in an afternoon, and your agent can do most of the work.

If you get stuck on any step, find me on X at [@timithechef](https://x.com/timithechef).</content:encoded><category>ai</category><category>agents</category><category>t3 code</category><category>vps</category><category>guide</category></item><item><title>How I Ship Without Reading Code</title><link>https://blog.timi.click/how-i-ship-without-reading-code/</link><guid isPermaLink="true">https://blog.timi.click/how-i-ship-without-reading-code/</guid><description>How to use SLADE—Skill-Led Agentic Development &amp; Engineering—to supervise agents and ship software.</description><pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate><content:encoded>I don’t read every line of code that my agents write.

That sounds irresponsible until you understand what I mean. I’m not merging code I have never thought about, and I’m not treating tests as magic. I supervise the intent, inspect the plan, keep implementation bounded, run tests, and make independent reviewers keep looking until the work is good enough to ship.

I call the method **SLADE**: Skill-Led Agentic Development &amp; Engineering. It is my way of using skills as process supervision, so agents can move from an initial idea to a reviewed and shipped pull request.

## From copy-paste to agents

I started the same way many people did. I would ask [ChatGPT](https://chatgpt.com/) or Google’s [Bard](https://en.wikipedia.org/wiki/Google_Bard) a question about some code, copy an entire file into the chat, explain what I wanted changed, then copy the answer back into my editor.

It was manual, but it gave me a glimpse of what coding could become. I couldn’t have predicted the exact shape of agentic development. I just knew the work was going to change.

[GitHub Copilot](https://github.com/features/copilot) removed some of the copying with tab completion. Based on the first few lines, it could suggest the next line, variable, function, or class. Then [Cursor](https://cursor.com/) made prompting part of the editor itself. I was writing fewer lines and more instructions.

Later, tools such as [Claude Code](https://www.anthropic.com/claude-code) and [Codex](https://github.com/openai/codex) made it natural to work with an agent that could inspect a repository, run commands, change files, commit code, and open a pull request. The editor became less important. Planning became more important.

In January 2026, I posted that I had reached the top one percent of Cursor users with three billion tokens processed. I use the screenshot as a usage marker; the engineering-quality evidence comes from the review and shipping loop. It captures how much of my work had moved into prompting and iteration.

&lt;blockquote class=&quot;twitter-tweet&quot;&gt;
  &lt;p lang=&quot;en&quot; dir=&quot;ltr&quot;&gt;Ik I&apos;m late to the party&lt;br&gt;&lt;br&gt;But being in the top 1% of Cursor users is wild. 3B tokens, damn &lt;a href=&quot;https://t.co/q1gxImVCMh&quot;&gt;https://t.co/q1gxImVCMh&lt;/a&gt;&lt;/p&gt;
  &amp;mdash; timi the chef 👨🏾‍🍳 (@timithechef) &lt;a href=&quot;https://twitter.com/timithechef/status/2008136073223029101?ref_src=twsrc%5Etfw&quot;&gt;January 5, 2026&lt;/a&gt;
&lt;/blockquote&gt;

Eventually, the output became too large for me to supervise line by line. I needed a way to supervise the process that produced the code.

## SLADE starts with intent

A detailed feature spec is useful. The starting point can also be a PRD from a product manager, a feature request from a user, a bug report, or a conversation with an agent that eventually becomes clear enough to act on.

The first skill is [**kickoff**](https://github.com/Timmyy3000/skills/blob/main/skills/kickoff/SKILL.md). It captures the intent, understands the repository context, chooses the planning mode, and routes the work to the next stage. It stays thin: planning, implementation, review, and delivery remain separate stages.

This general idea of portable, composable skills has also been formalized by [Anthropic’s Agent Skills](https://claude.com/blog/skills): folders that package instructions, scripts, and resources an agent can load when they are relevant. My skills apply that idea to the software-development lifecycle.

Once the intent is clear, the loop looks like this:

```text
Intent
  ↓
kickoff and context capture
  ↓
plan-it
  ↓
adversarial review
  ↓
simplicity review
  ↓
optional human approval
  ↓
ship-it
  ↓
reviewed PR
```

- **[plan-it](https://github.com/Timmyy3000/skills/blob/main/skills/plan-it/SKILL.md):** turns the intent into scope, non-scope, affected files, phases, acceptance criteria, validation, and risks.
- **[adversarial-review](https://github.com/Timmyy3000/skills/blob/main/skills/adversarial-review/SKILL.md):** attacks the plan for missing requirements, hidden dependencies, bad sequencing, and untested risks.
- **[simplicity-review](https://github.com/Timmyy3000/skills/blob/main/skills/simplicity-review/SKILL.md):** asks whether the same requirements and safeguards can be met with less machinery.
- **[ship-it](https://github.com/Timmyy3000/skills/blob/main/skills/ship-it/SKILL.md):** executes the accepted plan through implementation, validation, code review, and PR readiness.
- **[code-review](https://github.com/Timmyy3000/skills/blob/main/skills/code-review/SKILL.md):** gives the finished branch a fresh pass for correctness, regressions, security, and missing tests.
- **[create-pr](https://github.com/Timmyy3000/skills/blob/main/skills/create-pr/SKILL.md):** packages the result with a useful summary and test plan instead of producing another vague PR description.

For larger work, I approve the plan before implementation. I use [Lavish](https://github.com/Timmyy3000/lavish-axi), an HTML plan editor, to read the plan, comment on specific sections, and discuss changes with the agent. The gate is there to catch a bad direction before it becomes a large diff.

## Ship-it is where the code moves

After the plan is approved, [ship-it](https://github.com/Timmyy3000/skills/blob/main/skills/ship-it/SKILL.md) takes over. It breaks the work into bounded packets with explicit ownership, dependencies, acceptance criteria, and validation.

The implementation follows [Red–Green TDD](https://en.wikipedia.org/wiki/Test-driven_development) when practical:

1. **Red:** write a test that expresses the behavior and fails.
2. **Green:** implement the smallest change that makes it pass.
3. **Refactor:** improve the implementation without changing the contract.

This matters because agents that write tests after the implementation tend to write worse tests. They already know how the code works, so the tests often describe the implementation instead of defining the behavior. Writing the failing test first gives the work a contract before the agent starts looking for a way to make its code pass.

When several packets can move independently, [Forest](https://github.com/Timmyy3000/git-forest) gives each loop its own Git worktree. I built it to make parallel agent work practical: multiple features can move through the same codebase without fighting over one working directory or branch.

The loop also supports delegated implementation. `ship-it` keeps the current agent as the orchestrator, then lets it resolve implementation as `never`, `auto`, or `always`. When delegation is enabled, the orchestrator turns the plan into bounded packets with exclusive ownership, dispatches dependency-ready packets in parallel, integrates the results, and reruns the checks itself.

That changes the model economics. A stronger model can conduct the work, make the shared decisions, and review the result while a fleet of smaller, cheaper workers handles well-bounded packets. Independent work can finish sooner and cost less than asking one expensive model to do everything. Worker reports are only inputs; the orchestrator inspects their diffs and validates the integrated result.

## The loop in production

The best example is [Nabu PR #13](https://github.com/Timmyy3000/nabu/pull/13), which added agent-first temporary shared spaces. Nabu is a Markdown-native, agent-first knowledge OS. This feature lets an agent inspect a requested folder, show the user what will be shared, get confirmation, create a temporary live shared space, and generate a one-time invite another human or agent can redeem.

I had already done the product thinking. The feature request covered the goals, non-goals, workflows, security model, tests, acceptance criteria, and API surface. Kickoff moved that decided intent into the engineering loop.

The first run moved through the stages in order: plan approval, adversarial review, implementation, Red–Green TDD, validation, independent review, and PR handoff.

&lt;div class=&quot;workflow-gallery&quot; aria-label=&quot;Nabu kickoff workflow screenshots&quot;&gt;
  &lt;figure&gt;
    &lt;img src=&quot;/images/slade/nabu-plan-approval.png&quot; alt=&quot;Nabu kickoff showing the reviewed implementation plan before code changes begin&quot; /&gt;
    &lt;figcaption&gt;&lt;strong&gt;Plan approval.&lt;/strong&gt; The plan is reviewed before implementation begins.&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
    &lt;img src=&quot;/images/slade/nabu-adversarial-review.png&quot; alt=&quot;Nabu kickoff showing a fresh-context adversarial review of the plan&quot; /&gt;
    &lt;figcaption&gt;&lt;strong&gt;Adversarial review.&lt;/strong&gt; A fresh context attacks the plan before implementation gets expensive.&lt;/figcaption&gt;
  &lt;/figure&gt;
  &lt;figure&gt;
    &lt;img src=&quot;/images/slade/nabu-code-review.png&quot; alt=&quot;Nabu workflow showing an independent code review of the implementation branch&quot; /&gt;
    &lt;figcaption&gt;&lt;strong&gt;Independent review.&lt;/strong&gt; A separate pass checks the branch before the PR is handed off.&lt;/figcaption&gt;
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class=&quot;post-stats&quot; aria-label=&quot;Observed Nabu ship-it run details&quot;&gt;
  &lt;div&gt;&lt;strong&gt;66m 24s&lt;/strong&gt;&lt;span&gt;first ship-it run&lt;/span&gt;&lt;/div&gt;
  &lt;div&gt;&lt;strong&gt;175&lt;/strong&gt;&lt;span&gt;tests passed in the run&lt;/span&gt;&lt;/div&gt;
  &lt;div&gt;&lt;strong&gt;5/5&lt;/strong&gt;&lt;span&gt;later Enkii review state&lt;/span&gt;&lt;/div&gt;
&lt;/div&gt;

The PR was substantial, but I’m using it as workflow evidence rather than a file-by-file tour. The plan was reviewed, the implementation was bounded, the tests were written as part of the work, and the result passed through independent review before it merged and went to production.

## The PR keeps looping

After [create-pr](https://github.com/Timmyy3000/skills/blob/main/skills/create-pr/SKILL.md) opens the pull request, [Enkii](https://github.com/Timmyy3000/enkii) reviews it. Enkii is a tool I built for AI-powered pull-request review. It checks three things:

- code quality and bugs;
- security;
- repository-defined policy.

The agent checks for new review output, fixes the findings, and waits for another pass. The loop continues until every review lane reaches five out of five.

On Nabu PR #13, Enkii found a token-revocation mapping bug, an inconsistent duration default, and UI authorization paths that could turn an access failure into a 500 error. The security pass checked path traversal, symlink boundaries, scoped authorization, hashed secrets, atomic invite redemption, and revision-aware writes.

That is the difference between an agent that produces a diff and an engineering system that produces software I’m willing to ship.

## What I’m actually doing

I spend less time acting as a human syntax checker and more time deciding what should be built, whether the plan makes sense, whether the tradeoffs are acceptable, and whether the evidence is strong enough to ship.

That is how I can move tens of pull requests through a week, sometimes involving tens of thousands of lines of code. Those are my numbers, not a benchmark or a promise for everyone else.

&lt;div class=&quot;velocity-card&quot; aria-label=&quot;Engineering output over the last thirty days&quot;&gt;
  &lt;div class=&quot;velocity-card-top&quot;&gt;
    &lt;div&gt;
      &lt;span class=&quot;velocity-kicker&quot;&gt;ENGINEERING VELOCITY&lt;/span&gt;
      &lt;h3&gt;Last thirty days&lt;/h3&gt;
    &lt;/div&gt;
    &lt;div class=&quot;velocity-total&quot;&gt;&lt;strong&gt;68&lt;/strong&gt;&lt;span&gt;PRs merged&lt;/span&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class=&quot;velocity-metrics&quot;&gt;
    &lt;div&gt;&lt;strong&gt;10&lt;/strong&gt;&lt;span&gt;unique repos&lt;/span&gt;&lt;/div&gt;
    &lt;div&gt;&lt;strong&gt;496&lt;/strong&gt;&lt;span&gt;commits&lt;/span&gt;&lt;/div&gt;
    &lt;div&gt;&lt;strong&gt;+61.4k&lt;/strong&gt;&lt;span&gt;lines added&lt;/span&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class=&quot;velocity-table-wrap&quot;&gt;
    &lt;table&gt;
      &lt;caption&gt;Scope breakdown&lt;/caption&gt;
      &lt;thead&gt;&lt;tr&gt;&lt;th&gt;Scope&lt;/th&gt;&lt;th&gt;Repos&lt;/th&gt;&lt;th&gt;Commits&lt;/th&gt;&lt;th&gt;Code changes*&lt;/th&gt;&lt;th&gt;PRs opened&lt;/th&gt;&lt;th&gt;PRs merged&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;
      &lt;tbody&gt;
        &lt;tr&gt;&lt;td&gt;Docsyde&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;363&lt;/td&gt;&lt;td&gt;+45,010 / −8,669&lt;/td&gt;&lt;td&gt;45&lt;/td&gt;&lt;td&gt;47&lt;/td&gt;&lt;/tr&gt;
        &lt;tr&gt;&lt;td&gt;Open source / agent-side repos&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;131&lt;/td&gt;&lt;td&gt;+16,363 / −10,421&lt;/td&gt;&lt;td&gt;19&lt;/td&gt;&lt;td&gt;17&lt;/td&gt;&lt;/tr&gt;
        &lt;tr class=&quot;combined&quot;&gt;&lt;td&gt;Combined&lt;/td&gt;&lt;td&gt;10 unique&lt;/td&gt;&lt;td&gt;496&lt;/td&gt;&lt;td&gt;+61,412 / −19,098&lt;/td&gt;&lt;td&gt;68&lt;/td&gt;&lt;td&gt;68&lt;/td&gt;&lt;/tr&gt;
      &lt;/tbody&gt;
    &lt;/table&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;small&gt;* Lines added and removed in the report.&lt;/small&gt;

SLADE is what makes this output sustainable. Agents handle more of the implementation; the system supplies intent, plans, bounded packets, tests, independent review, and explicit shipping gates. That is how I can move quickly without turning speed into guesswork.

Other approaches can work. Some people prefer multi-agent orchestration, loop engineering, graph engineering, or a much simpler single-agent workflow. SLADE is the arrangement I found useful because it lets me give agents more responsibility without giving up process control.

I supervise the intent, the plan, the boundaries, the tests, the reviews, the policies, and the final evidence. The agents handle more of the implementation, but they do it inside a system that keeps asking whether the work is correct, simple enough, secure enough, and ready to ship.

That is what I mean when I say I ship without reading code.

## Get started

Tell your agent to install the [kickoff skill](https://github.com/Timmyy3000/skills/blob/main/skills/kickoff/SKILL.md):

```bash
npx skills add Timmyy3000/skills --skill kickoff
```

The skill has enough context to get you started. Good luck building your own SLADE loops.</content:encoded><category>ai</category><category>agents</category><category>software engineering</category><category>SLADE</category></item><item><title>The Mind–Body Problem of AI</title><link>https://blog.timi.click/mind-body-problem-of-ai/</link><guid isPermaLink="true">https://blog.timi.click/mind-body-problem-of-ai/</guid><description>LLMs, harnesses, and agents—and how to become a first-class citizen in the age of agents.</description><pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate><content:encoded>A friend asked: “Is Codex an LLM or an agent? What is OpenClaw? And what exactly is the difference between all of this?”

The confusion is understandable. AI products often bundle several layers under one name, then expect everyone else to understand the difference.

Here’s my attempt to make the landscape less confusing: I’m borrowing a model from something we all already understand—a human being.

## The mind

Start with the [LLM](https://en.wikipedia.org/wiki/Large_language_model).

An LLM is like a mind inside a computer. It receives an input, thinks through patterns it has learned, and produces an output. The quality of that output depends on what the model learned during training, how it was tuned, and what information it receives in the moment.

If you want a sense of what these minds look like in practice, look at the current model catalogues from the major labs: OpenAI’s [GPT-5.6 Sol](https://developers.openai.com/api/docs/models), Anthropic’s [Claude Opus 5](https://platform.claude.com/docs/en/about-claude/models/overview), Google’s [Gemini models](https://ai.google.dev/gemini-api/docs/models), and Meta’s [Llama 4](https://developer.meta.com/ai/models/llama-4/). The open-model side includes releases such as [Kimi K3](https://platform.kimi.ai/docs/overview), [DeepSeek V4](https://api-docs.deepseek.com/), and [GLM 5.2](https://blogs.nvidia.com/blog/open-secure-ai-alliance/). Put the same prompt into each one and you will get very different responses. They differ in how they reason, what they can do, what it costs to access them, and where they fail.

There is a thought experiment that helps here: the [brain in a vat](https://en.wikipedia.org/wiki/Brain_in_a_vat). A brain receives signals that stand in for an entire world, but has no body with which to touch that world. An LLM is something like that: a mind that can process what reaches it, but cannot act on anything by itself.

A mind can solve a problem in its head. It can imagine lifting a table. It can explain how to build a house. But imagining the action is not the same as doing it.

A bare model can produce text, code, or a structured response. It does not independently open your filesystem, run a command, browse the web, or change the world outside the computation that produced its output.

That distinction matters because a model can be extremely capable and still be unable to accomplish a task on its own.

## The body

The harness is the body.

Give a mind a hand and it can lift something. Give it eyes and it can see. Give it a mouth and it can speak to another person. The body is the mechanism through which the mind interacts with the outside world.

A harness does something similar for an LLM. It gives the model ways to act:

- **Tools:** Read files, run shell commands, open applications, browse the web, play games, call APIs, or interact with a product environment.
- **Context:** Provide the codebase, documents, conversation history, or other information needed for the task.
- **Permissions:** Decide what the system is allowed to read, change, send, or execute.
- **Interface:** Give a person a way to direct the system: a terminal, desktop app, chat, API, or something else.

A [command-line interface](https://developer.mozilla.org/en-US/docs/Learn_web_development/Getting_started/Environment_setup/Command_line) is one kind of body surface. It is a text interface for executing programs. For example, the [Codex CLI](https://github.com/openai/codex) is a coding harness delivered through that surface: the terminal is how you interact with it; the model is the mind doing the reasoning; the surrounding software is the body that lets it work on a codebase.

## The agent

An agent is what you get when the mind can act through the body.

&gt; **Mind + body = an agent that can reason and act.**

The industry uses “agent” loosely. Sometimes it means a model with a tool. Sometimes it means a system that works through a complex task for hours. For this article, the useful distinction is simpler: an agent is a mind with a body that can act.

A human being can take in the world, think through a problem, and use a body to do something about it. An AI agent does a similar thing through software: it receives a task, reasons about what to do, uses the tools available to it, observes what happened, and continues.

The system may also have internal state, objectives, loops, or policies. Those are important for how it operates. They are not the central point of this analogy. The central point is simpler: the model supplies the reasoning, and the harness gives that reasoning a way to act.

## What the mind learns along the way

Human beings are not defined only by having a mind and a body. What we learn, remember, notice, and instinctively reach for changes what we can do with them.

A person who has learned how to use a saw is more useful in a workshop than someone who has never seen one. A doctor carries specialized knowledge. A person remembers what went wrong last time. The subconscious notices patterns before the conscious mind can explain them. Instinct can move a person before deliberation catches up.

AI systems have rough equivalents:

- **Skills:** Reusable instructions, tools, code, or workflows for a particular kind of work.
- **Memory:** Information carried across interactions or retrieved when it becomes relevant.
- **Specialization:** Design choices that make the system better at one kind of job than another.

These things make an agent more capable. They are not necessarily the core model, and they are not what makes the system an agent in the first place. They are closer to the things that shape a person over time: knowledge, memory, instinct, practice, and the environments in which they have learned to operate.

## Freaky Friday: one mind, different bodies

This is where the analogy breaks, but in a useful way!

Humans cannot swap minds and bodies. We do not get to put one person’s brain into the body of the world’s best swimmer and see what happens. We certainly do not get a *Freaky Friday* menu where we choose a new body for the same mind.

AI systems can be assembled more like that.

If you like how a particular model reasons, you are not automatically restricted to the one product that first exposed it to you. Depending on compatibility, licensing, cost, and access, the same model can work through different harnesses.

One body may be built for software engineering. Another may be built for security work. Another may be built to manage a person’s messages, files, calendar, and long-running responsibilities. The mind may be similar; the resulting agent can feel completely different because it has a different body and a different job.

That is also why these systems feel so different even when they use models with similar abilities.

[Codex CLI](https://github.com/openai/codex) gives a model a body shaped around software engineering. [OpenCode](https://opencode.ai/docs/) gives it another coding-oriented body, with its own providers, permissions, tools, and extension points.

[Hermes](https://hermes-agent.nousresearch.com/docs) and [OpenClaw](https://github.com/openclaw/openclaw) are shaped more like general personal assistants: they are built to move across conversations, files, reminders, channels, and long-running responsibilities.

My own product, [Docsyde](https://usedocsyde.com), gives an agent a body designed for document work—drawing context from documents and CRM systems rather than treating a codebase as its natural habitat.

## How to choose your stack

Once you separate the layers, choosing an AI system becomes less like picking a mascot and more like building a cyborg. You are weighing different minds, bodies, tools, costs, and behaviours, then deciding which combination fits the way you want to work.

- **Budget:** The most capable option is useless if its price or usage limits make it impractical.
- **Model:** Different labs and models have different habits, strengths, weaknesses, and styles of reasoning.
- **Reasoning:** Many products let you trade speed and cost for more deliberate reasoning.
- **Harness:** What can it read, write, execute, browse, remember, or connect to?
- **Job:** Do you need a coding agent, a research assistant, a security tool, or a personal assistant?
- **Openness:** Can you inspect, modify, self-host, or replace the parts that matter to you?

One resource I find useful for figuring out how models stack up—and how they perform across different harnesses—is [Artificial Analysis](https://artificialanalysis.ai/). It helps you compare models by capability, speed, and pricing instead of trusting somebody’s permanent declaration that one model is “the best.”

## Open source as biohacking

Open source makes the human analogy stranger—and more exciting.

As human beings, we are mostly “closed source,” if you think about it. We arrive with bodies we did not design and minds shaped by biology, upbringing, and experience. We can modify ourselves, but only within limits set by what we inherited.

Open source is closer to biohacking. It gives people access to the underlying machinery. They can inspect the system, understand how it works, change parts of it, and build a new version for a different purpose.

That can happen at different layers. An open-weight model gives you access to learned parameters that can run outside one hosted product. An open-source harness gives you code you can inspect and modify. A genuinely open-source AI system makes broader freedoms available across the components needed to study, change, and share it.

Those distinctions matter. “Open” is not one switch. According to the [Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition), an open-source AI system should let people use, study, modify, and share it. The preferred form for modifying a model includes information about its training data, the complete source code used to train and run it, and the model parameters. A model can be open-weight without meeting that definition. A product can expose an open-source CLI while keeping other parts of the overall service closed.

And this is no longer a fringe conversation. NVIDIA’s [Open Secure AI Alliance](https://blogs.nvidia.com/blog/open-secure-ai-alliance/) brings together companies including NVIDIA, Hugging Face, Mistral, Microsoft, GitHub, Cloudflare, Perplexity, OpenClaw, Nous Research, Thinking Machines Lab, and many others to build and share open tools for AI safety and security. Its argument is practical: defenders need systems they can inspect, adapt, and run on their own infrastructure.

The labs driving the open-weight frontier are not all in one country. Meta’s Llama, OpenAI’s GPT-OSS, DeepSeek’s open releases, and Kimi’s model family all show different versions of the same pressure: people want capable models they can run, study, connect to their own harnesses, and improve. The details differ by release and license, but the direction is unmistakable.

The more of the mind and body people can inspect and change, the more freedom they have to combine systems for their own purposes.

Understanding the parts means you can make better decisions about the systems you use. You can ask what kind of mind you want, what kind of body it needs, and what agent the two can create together.</content:encoded><category>ai</category><category>agents</category><category>models</category><category>open source</category></item><item><title>Hello World!</title><link>https://blog.timi.click/hello-world/</link><guid isPermaLink="true">https://blog.timi.click/hello-world/</guid><description>My first real stab at a serious blog — journaling experiments, AI agents, and building in public.</description><pubDate>Thu, 12 Feb 2026 00:00:00 GMT</pubDate><content:encoded>Of course I use AI to write, get off my back.  

I am not really much of a writer, but this is a humble attempt to communicate my thoughts to the world, to people, and to the LLMs that will inevitably end up crawling this blog someday.

This is my first serious stab at blogging, and I plan to use it as a journal — a place to document my experiences, experiments, and the things I&apos;m learning as I build.

## What I&apos;m focused on

My biggest interest right now is **AI agents**. I&apos;m building an AI company called [Docsyde](https://x.com/docsyde_ai), and I&apos;m deep in the trenches of figuring out what autonomous agents can actually do when you let them loose on real problems. If you want to follow along or say hi, you can find me on X at [@timithechef](https://x.com/timithechef).

## I&apos;ve been a founder before

Before Docsyde, I built a fintech company. It didn&apos;t last forever, but the lessons I took from that experience are ones I wouldn&apos;t trade for anything. There&apos;s a kind of knowledge you only get from actually doing the thing — from shipping, from failing, from talking to users who don&apos;t care about your roadmap. That experience shaped how I think about building, and it&apos;s a big part of why I&apos;m writing now.

## About this website

Here&apos;s something worth mentioning: **this entire site was built by Claude and Codex**. I didn&apos;t write a single line. Not the layout, not the styling, not the theme switcher — all of it was shipped by AI agents.

Speaking of which — try the theme switcher. I told my agents to add it, and they came up with 20 different themes. It&apos;s a small thing, but it&apos;s a good example of how creative AI agents can get when you give them room.

I&apos;m committing to writing **zero code** on this project. The only thing I&apos;ll write by hand are the articles themselves. Everything else — every feature, every fix, every design decision — gets delegated to agents. I want this site to be living proof of what AI can ship, so that people who are still skeptical can come here, poke around, and see for themselves.

## What&apos;s coming

Expect posts about AI agents, autonomous workflows, lessons from building startups, and whatever experiments I&apos;m running at any given time. Some of it will be polished. A lot of it won&apos;t be. That&apos;s the point — this is a journal, not a magazine.

Let&apos;s see where this goes.</content:encoded><category>thoughts</category><category>ai</category></item></channel></rss>