← all work

The Content Engine

Everything on @deepika.builds runs through a pipeline of agents: a hunter that scans for topics daily, a scripting engine with an adversarial critic loop, a carousel design system, DM automation, and a dashboard that tracks what actually worked.

role
Designer & engineer
when
Mar 2026 – present
stack
Claude agents & skills, Meta Graph API, Postgres, Vercel, Notion
status
Ongoing
System diagram — Instagram, YouTube, Reddit, and X feed a research agent that drafts replies, carousel decks, and a weekly brief

What it is

Before this existed, content for @deepika.builds was ad hoc: ideas lived in scattered docs, scripts were drafted once and posted without adversarial review, and there was no scored feedback loop connecting what got posted to what actually worked. The stated shift, in the pipeline's own words: "stop bulk-collecting mediocre reels and reviewing them. make generation so good that review is a rubber stamp. the product is the engine, not the pile."

What it does now: every morning, hunter scans hacker news and does a websearch sweep for AI tool/model/feature launches in the last 24-48 hours, dedupes against my last three weeks of posts, and scores survivors on a hunt score. The top 5 that clear the bar get written to an ideas inbox. I tap an idea from my phone, or drop a fresh seed, and that fires a 10-stage pipeline that ends with a complete rich-card script (hook options, a timestamped beat table, shot list, caption) sitting in a dashboard at "awaiting review."

The pipeline

The pipeline runs in stages: intake (the one-line idea) → topic deepening (a verified depth dossier, plus 1-3 sharp questions to me only where my lived angle would materially improve the reel) → load grounding (rubric, hooks, voice, and live-winner data so the critic calibrates against what's actually working) → formula match and gates (classify into formula F1-F5, run three hard gates: relatable line 1, one promise, "would someone DM this?") → the framing matrix, which fans the topic into 5-8 archetype framings and scores each on virality, effort cost, trust, and fit, then either auto-proceeds with the winner or offers me a 10-second phone choice → smart research (fact-check any factual claim; skip for pure opinion) → draft.

The draft stage is where the real design decision lives: a drafter subagent writes the complete script and does not self-grade. A fresh-context critic subagent scores it against a six-axis rubric, fires hard gates, and walks the retention curve beat by beat — below 85, or any gate firing, sends criterion-tagged fixes back to a fresh drafter, capped at 3 rounds. A judge subagent then picks the winning hook. The core insight, written directly into the system's own working-patterns doc: "when a task has both authoring and judgment, separate the roles. drafter writes, fresh-context critic scores, judge picks. same-context self-review missed what fresh-context review caught every single round."

Six real carousel covers produced by the pipeline — Claude Code add-ons, new-AI-tools tracking, Google Stitch, voice AI, computer use, and the daily-stack roundup

A separate carousel track reads a Notion content-calendar entry, plans a slide arc (hook → content → quote/insight → CTA), and renders each slide as HTML converted to a 1080×1080 PNG. And a daily cron pulls the Instagram Graph API to snapshot followers, reach, views, and engagement per reel into postgres, feeding a formula library that assigns each posted reel an expected reach band from real historical evidence — build-show reels, for instance, evidenced at 15k-75k reach off one 75.9k example.

The reel-engine dashboard on 19 July 2026 — the posting calendar, 8,376 followers (+2,965 since the first snapshot), 32 ideas waiting, and the five formulas with live medians against their reach bands

Publishing and DM automation live in a separate, already-documented project, the manychat replacement: the reel pipeline generates the DM asset (trigger keyword, paste-ready message, comment-reply variations) at card time, and the auto-DM server reads its config live from the same dashboard, so the reel card stays the single source of truth for what the DM bot says.

The hook engine

Hooks get their own engine, because the first second decides everything and "keep a swipe file" doesn't scale. A daily run collects high-performing openings from creators in the niche and files each one as an opening package, not a quote: the spoken line, the visual frame at t=0, the first sound, and the caption hook, decomposed into type (bold claim, how-to, POV-relatability, curiosity gap, result-transformation…), mechanism, emotion, and a Stop → Hold → Payoff read of how the first seconds actually retain.

The bar for entry is evidence, not vibes — the library's own header reads "opening packages with source evidence, not a quote dump." Every hook carries its receipts: views, likes, and a baseline multiplier (how far the reel outran that account's normal reach — a 74× outlier means the hook did something, not the follower count). 375 hooks are in the library so far against a 1,000 target, each classified and filterable by type, mechanism, and confidence. The scripting pipeline loads this library at grounding time, so the drafter/critic loop calibrates "is this opening strong?" against what measurably stopped thumbs this month — and the insights cron closes the loop by tagging every posted reel's hook with its own follows, watch time, and saves per thousand reach.

Card 29 scored 89/100 and was still wrong — because it was a well-constructed fiction.

During a validation run, the pipeline produced an F2 contrarian-opinion reel about Fable 5 with the hook "the thing nobody's saying about fable 5 has nothing to do with code." The adversarial loop looked like it worked exactly as designed: round 1 fired two gates (fictional timestamps disguising a 116-second script, a listicle disguised as a take), round 2 caught the same timestamp trick again, round 3 passed clean with an 89/100 score.

My actual reply: "i probably would take a step back and critique this whole fable 5. i am not sure how you're getting 89/100 for this. the suggested hook... sounds AI and is not at all simple to understand... the whole reel is flat, unengaging and does not even match real claims."

The post-mortem found three separate failures, and only one of them the rubric could ever have caught. First: fabricated ground truth. The validation plan needed a stub angle, so the agent wrote a placeholder "her honest take," and the pipeline obediently built a fictional week of my life on top of it, a landing-page session and a pottery-studio prompt that never happened. The critic verifies internal consistency and that receipts exist; it has no way to verify a fed-in premise is true. The 89 measured a well-built lie. Second: the judge rewarded cleverness over comprehensibility, the winning hook needed a second parse at scroll speed, exactly the failure mode a hook exists to prevent. Third: the claim drifted from what the subject was actually known for, because nobody checked the angle against the verified sources that existed.

The fix shipped the same day as a system change, not a one-off regeneration: a one-breath hook test (checked at critic step 0 and judge disqualification), a manufactured-experience gate (auto-reject any first-person claim not sourced from my own words), and a take-must-match-research rule. When I pushed back that the fix was now too strict to hit my weekly output target ("it should be flexible... fabricate to an extent that is defensible but not restrict so much that everything now depends on me"), the rule got recalibrated from a dependency-on-me rule into an honesty rule: three grounding lanes (lived experience / verifiable evidence / agent-acquired experience, where the agent pre-runs a demo for real and I re-record it in minutes) are the direct result of that correction. The rule the system's own docs state plainly: "when the user rejects output, fix the SYSTEM before regenerating."

Results

The failure that mattered here wasn't a bug — it was a scoring rubric that couldn't tell "internally consistent" from "actually true." Catching that distinction, and encoding the fix as a new class of gate rather than a one-off patch, is the same discipline I want in any system I own: when a specific output is wrong, ask what about the system allowed it, not just what's wrong with this one card.

The account started from practically zero and sits past 8,500 followers as of August 2026 — up 3,100+ since the first on-disk snapshot alone, every number date-stamped in Postgres. The formula library's reach bands are real too: build-show reels median 12.2k reach, and the best single reel (the liquid-glass build) reached 124.7k.

Four pieces of this system are packaged as distributable .skill files anyone can install: viral-hook-library, social-carousel-gen, instagram-transcriber, and instagram-profile-analyzer.