← home

feeling the vibe

vibe coding is back in style!


vibe coding is so back. except now we call it agentic engineering. and the code is so much better than it used to be. except for when it isn’t. then the vibes start to fall short.

the last time i was full on vibe coding was after Claude Code had just come out, and i was putting claudius through the GaLt torture chamber. the code was “cheap”, and the quality was good enough to have something working after an hour or so. it struggled from the same problem as vibe coding did before, through tools like cursor, which was that the codebases became unworkable, and disgusting to trudge through, as well as poor expandablity.

these days, i find myself using Pi, which markets itself with “There are many agent harnesses, but this one is yours”. Pi’s greatest strength is its ability to write for itself, allowing it to expand and modify to fit the needs of any individual. no two installs of Pi will be the same! for example, mine looks like this on startup:

my pi

i will admit that it is easy to give yourself a lot of bloat initially when first trying it out, but once you figure out what it is exactly that you need, it becomes extremely minimal, which is what Pi excels at. anytime i need it to do something, i can ask it to make itself a skill or tool, and it will do just that. i once asked it to build itself an extension to search the web using a self hosted SearxNG instance. it took a few minutes, (and some cheap inference courtsey of my Opencode Go subscription), and voila! a working web search tool for my Pi to use whenever!

export default function (pi: ExtensionAPI) {
  pi.registerTool({
    name: "web_search",
    label: "Web Search",
    description: "Search the web via the local SearXNG instance (privacy-respecting metasearch).",
    promptSnippet: "Search the web for current or external information using SearXNG",
    promptGuidelines: [
      "Use web_search when the user asks for web/current/external information.",
      "Use web_search before guessing when facts are likely outside the repo.",
    ],
    parameters,

    async execute(_toolCallId, params: WebSearchParams, signal) {
      const searxngUrl = (process.env.SEARXNG_URL || DEFAULT_SEARXNG_URL).replace(/\/+$/, "");
      const timeoutSeconds = Math.min(params.timeout ?? DEFAULT_TIMEOUT_SECONDS, MAX_TIMEOUT_SECONDS);
      const controller = new AbortController();
      const timeoutHandle = setTimeout(
        () => controller.abort(new Error("SearXNG request timed out")),
        timeoutSeconds * 1000,
      );
      const abortForwarder = () => controller.abort(signal?.reason ?? new Error("Request cancelled"));
      signal?.addEventListener("abort", abortForwarder, { once: true });

      try {
        const workflow = params.workflow ?? "search";
        const limit = params.limit ?? DEFAULT_LIMIT;
        const category = workflowToCategory(workflow);

        const searchParams = new URLSearchParams({
          q: params.query,
          format: "json",
          safesearch: params.safeSearch === false ? "0" : "1",
        });
        if (category) searchParams.set("categories", category);
        if (typeof params.page === "number") searchParams.set("pageno", String(params.page));

        const url = `${searxngUrl}/search?${searchParams.toString()}`;

        const response = await fetch(url, {
          method: "GET",
          headers: {
            "Accept": "application/json",
          },
          signal: controller.signal,
        });

        const raw = await response.text();
        if (!response.ok) {
          throw new Error(`SearXNG error ${response.status}: ${truncate(raw, 600)}`);
        }

        let json: any;
        try {
          json = JSON.parse(raw);
        } catch {
          throw new Error("SearXNG returned non-JSON response");
        }

        const results = dedupeByUrl(collectResults(json?.results)).filter(
          (r) => r.title || r.url || r.snippet,
        );
        const picked = results.slice(0, limit);

        const infoboxTitles = Array.isArray(json?.infoboxes)
          ? json.infoboxes.map((ib: any) => firstString(ib?.title, ib?.name, "")).filter(Boolean)
          : [];
        const answers = Array.isArray(json?.answers) ? json.answers : [];
        const suggestions = Array.isArray(json?.suggestions) ? json.suggestions : [];

        if (picked.length === 0 && infoboxTitles.length === 0 && answers.length === 0) {
          return {
            content: [{ type: "text", text: `No results for: ${params.query}` }],
            details: {
              query: params.query,
              workflow,
              category,
              count: 0,
            },
          };
        }

        const lines: string[] = [];
        lines.push(`Results for: ${params.query}`);
        lines.push("");

        if (infoboxTitles.length > 0) {
          lines.push(`Infobox: ${infoboxTitles.join(" | ")}`);
          lines.push("");
        }

        for (let i = 0; i < picked.length; i++) {
          const r = picked[i];
          lines.push(`${i + 1}. ${r.title || "Untitled"}`);
          if (r.url) lines.push(`   ${r.url}`);
          if (r.snippet) lines.push(`   ${truncate(oneLine(r.snippet), 280)}`);
          lines.push("");
        }

        if (suggestions.length > 0) {
          lines.push(`Suggestions: ${suggestions.join(", ")}`);
        }

        return {
          content: [{ type: "text", text: lines.join("\n").trim() }],
          details: {
            query: params.query,
            workflow,
            category,
            count: picked.length,
            answers,
            results: picked,
          },
        };
      } finally {
        clearTimeout(timeoutHandle);
        signal?.removeEventListener("abort", abortForwarder);
      }
    },
  });
}

full version can be found over at this gist here.

well, i guess this means software is solved! anthropic has finally got it right! i mean it literally one shot an extension for itself with an absolutely terrible prompt! we will never need software engineers ever again!

…

WRONG

the newest frontier models often are more sycophantic than before, which is mostly visible in their behaviour towards users requesting features that are poorly (or not even) designed, or if the user doesn’t understand what they actually are trying to solve at all. the models will instead use every trick they have available to them to present to the user that what they requested, is indeed implemented, even if poorly, and often at times, purely visual and not functional. they can get away with this behaviour as they know that for many users, the full final output of text summarising their work, papercuts, and usually the git status, none of that will get read.

a vague prompt, such as:

make a version of spotify wrapped for Pi sessions. should go from jan 1st of the current year, and the most recent date (for now, as we approach the end of the year, we will change it to be 30th dec)

will then just result in the model trying to do its absolute best to produce exactly what i asked it to do, and the end result was indeed roughly what i asked for and expected, but it still isnt what i wanted.

initial pi wrapped iteration

i could keep iterating through further prompts (which, based on the first prompt, would likely stay vague enough for the models to fill in the many, many gaps), or i could start from scratch, and actually plan exactly what i want it to do, and look like.

so, in order to get exactly what i want, lets look at everything i didn’t like:

  • it created a node package that the user would have to download/clone first, and then run. ideally, the user runs it through node’s ‘npx’ as a one-off (or as many times as they like). the npm cli is authenticated, and ready to be used to publish new packages, which the model can do by itself.
  • it generated a html artifact only, whereas i would want the option of an image, with html as the default
  • it only estimated tokens for messages, not code, commands, tool calls etc.
  • it estimated the tokens, instead of using something like tiktoken to count them more accuratly
  • it designed the html to use theming similar to Spotify, rather than my usual design language that i keep in a skill file

using those, i can now give the model an actually usable prompt:

Build and publish an npm CLI called `pi-wrapped` that creates a “Pi Wrapped” report from my local Pi session history.

Read the relevant Pi documentation and inspect the local session format before you design the solution. Do not guess the schema.

## User experience

The package must work as a one-off command:

    npx pi-wrapped

The user must not need to clone a repository or manually install dependencies.

The npm CLI is already authenticated. After you test the package, publish it to npm yourself. 

## Date range

Include sessions from January 1 of the current year through the most recent date present in the session data.

Put the date-range logic in one clearly named place so it is easy to change the end date to December 30 later.

Show the exact date range in the report.

## Output

Generate an HTML report by default.

Also support image output:

    npx pi-wrapped --format image

Support an explicit output path:

    npx pi-wrapped --output ./pi-wrapped.html

Choose PNG as the image format unless the available rendering tools make another format clearly better.

The command must not upload session data. All processing must happen locally.

## Statistics

Calculate useful yearly statistics, including:

- number of sessions
- active days
- total user messages
- total assistant messages
- model usage
- tool calls, grouped by tool
- shell commands
- code and file-edit activity
- estimated token totals
- the longest or most active sessions
- notable usage patterns

Token totals must include all available session content, not only chat messages. This includes message text, code, commands, tool-call arguments, and tool results where they are stored.

Use `tiktoken`, or a compatible accurate tokenizer, instead of estimating tokens from character counts.

If the exact tokenizer for a model is not available, use the closest documented encoding and clearly label the result as an approximation. Document exactly which records and fields are included
in the token count.

Handle missing, malformed, or partially recorded data without crashing.

## Design

Do not copy Spotify’s visual style or branding.

Use the design language defined in `product-ui-craft`. Read and follow that skill before implementing the report.

The result should feel like a polished personal annual review. It must remain readable on desktop and mobile and when exported as an image.

## Implementation requirements

- Use Node.js.
- Provide a proper executable through the package’s `bin` field.
- Keep installation and runtime dependencies as small as practical.
- Detect the standard Pi session location automatically.
- Add a CLI option for a custom session-data path.
- Provide useful `--help` output.
- Do not expose private session content in the report by default.
- Add automated tests for parsing, date filtering, token counting, and output generation.
- Test the packed npm artifact with `npm pack` and execute that artifact before publishing.
- Write a concise README with command examples, data-source details, privacy behavior, and token-counting limitations.
- Before implementation, check whether `pi-wrapped` is available on npm.
   - If the package name is already taken:
     - Do not select an alternative name.
     - Do not publish any package.
     - Record the conflict with the `papercuts` tool.
     - Stop before publication and tell me that the requested name is unavailable.
     - Include the existing npm package URL and the papercut record in your final response.
   - Only publish when the exact requested package name is available.

## Acceptance criteria

The work is complete only when:

1. `npx pi-wrapped` generates a valid HTML report.
2. `npx pi-wrapped --format image` generates a valid image.
3. Only sessions in the specified date range are included.
4. Token counts include every supported content category.
5. The output follows my design skill rather than Spotify branding.
6. Tests and linting pass.
7. The packed package works in a clean temporary directory.
8. If the exact requested package name is available, publish the package to npm.
9. If the name is unavailable, do not publish. Record the conflict with the `papercuts` tool and report it to me.
10. Give me a final completion report containing:
   - whether the requested package name was available
   - the published package name and version, if publication occurred
   - the exact `npx` command, if publication occurred
   - the existing npm package URL, if the name was unavailable
   - generated file paths
   - test, lint, packaging, and publication results

that’s a lot more tokens that we initially sent! but it is worth it, as what it created this time is much closer to what i wanted. it wasn’t perfect, and it tripped itself up twice (one on purpose, the other not so much), but it did a great job overall.

this is how the session stats looked:

InputOutputReasoningCache hitCostContext
66k109k9.4M99.8%$0.10414.8%/1.0M

and the final product: final pi wrapped iteration

the final prompt isn’t perfect, for example, the statistics section is still vague, and only features a list of what should be included, rather than what exactly we want to show in the statistics section, leaving the model to fill in the gap itself. that is how everyone should be using these models, being responsible and conscious for everything you are making. no gaps! otherwise it just becomes the slop fest! and god knows the slop fest is not fun.

this post will be updated to include the Pi session transcripts at some point in the near future