Arkus Innovation Studios
Ship Log 002 · The week we stopped adding AI and started bounding it
Five repos, one pattern this week. Stop adding new AI surfaces. Start making the existing ones bounded, observable, and gated enough for a real operator.

This week the pattern across five repos was the same. Stop adding new AI surfaces. Start making the existing ones bounded, observable, and gated enough for a real operator. Single topic, five mini-logs underneath it.
Onward.
Arkus Suite: a federation, not a bundle
For most of the Suite's life, some product seams still depended on assumptions that needed to become explicit access checks. That works fine until you have real users with real entitlements, and then it stops working.
This week the Suite got an actual federation spine. The source-lock convention is documented. Suite JWTs are minted. Product guards now enforce access at the route boundary, not in scattered service code. Cipher artifact reads moved behind a server-only path that verifies Suite product access for the calling token, rejects local-preview subjects, and returns no-store cache headers.
Index got staged contract responses. The API routes for analyze, source upload, PDF, and the worker cron now enforce same-origin, authenticated user, entitlement, and persisted-user checks before returning a deliberately staged response. The workers, Supabase Storage writes, PDF generation, and provider calls stay disabled until the queue, timeout, retry, lease, and provenance behavior have been designed instead of assumed.
The restraint is the point. A federation is what you have when access, entitlements, and execution are governed. A bundle is what you have when products share a logo.
Cipher: evidence over performance
Cipher had been producing confident-looking output without making the evidence behind it easy enough to inspect. One real report showed the problem clearly: the score said 78, but the evidence behind that number was not easy enough to inspect.
This week the report became more defensible. The 0 to 100 score is now a true 0 to 100. The previous formula compressed weighted averages into the 20 to 100 band, which is the kind of thing that looks like a rounding choice and is actually a credibility problem. Page 9 of the report became Evidence Calibration, with score-moving distinctions, source posture, and explicit threshold versioning. Concept-based dedupe replaced phrase matching, so a report no longer says the same thing four ways because four different paragraphs framed it differently.
Extraction got reliability work. PDF extraction moved from long synchronous requests to async job polling. Jobs persist in an extract_jobs table, move through processing, completed, and failed, and surface stale jobs after 6 minutes. The user gets a durable job ID instead of a fragile open connection. This was a real production failure mode, not a hypothetical one.
Profile-aware evidence depth landed. The system now classifies a document as a long dense brief, a sparse visual deck, or a metric-dense deck, and budgets evidence extraction accordingly. A 60-slide deck with three real metrics no longer gets the same extraction depth as a 30-page operating memo.
Magic-link sign-in shipped for approved users, with localized strings for English, Portuguese, and Japanese. The Neon Auth handoff to the app domain is wired through a one-time bridge token, so the session is created on the right host.
The honest part: quota pressure made model routing an operating concern, not just a config choice. A rollback path was explored in a closed PR, but the larger lesson is that provider billing behavior needs a runbook. If you are routing across model tiers, write the runbook before you need it.
Prism/Index: green should mean green
Prism had a CI workflow that ran a hardcoded subset of tests, including at least one file that had been deleted. The suite was passing while hundreds of tests were not gating merges. This is the kind of slow-bleed problem that does not announce itself until you go looking.
The full Vitest suite now runs on every PR. 37 files, 711 tests, about 103 seconds. "Green" now means the whole repo.
Deep-research proxy endpoints now require auth. Per-user sliding-window rate limits sit in front of key paid model endpoints, backed by Upstash Redis. The Gemini proxy generationConfig got an allowlist, so future expensive parameters cannot be forwarded by accident. A QA gate that had been parsing CPC patent codes as impossible USD currency values got bounded with a proper regex.
Financial and market provenance caps were corrected. The previous logic was treating successful evidence-backed runs as unverified by default, which is exactly the wrong direction for a diligence product. Binding cap metadata now distinguishes the rule that actually moved the score from supporting candidates.
The EPO patent routes needed an unusual fix. Vercel cold-start Node File Trace was missing cross-directory imports from api/ into src/, so the OAuth token helpers got inlined into the route files. Not elegant. Worked.
VentureIP: from launch surface to operating surface
VentureIP moved from "looks ready" to "operates correctly."
The PRD got reframed. The product is an Evidence Workspace, not a theatrical terminal with fake verification QR codes and simulated logs. Route hygiene followed. Auth-adjacent routes moved to noindex, the sitemap stopped exposing redeem and request-access paths, the honeypot field got renamed away from the obvious "website," and admin pages got force-dynamic where they read live state.
The portal stopped leaning on static preview data. Fund briefs now parse from real content. The dashboard computes KPIs against real state. The activity feed reflects real activity. The Calendly URL prefills from real member context.
Admin operations got depth. Pagination on every list. Typed destructive-action confirmations like RESCIND, SUSPEND, and DISABLE. The audit, member, lead, and operational event surfaces all move at production volume instead of preview volume.
CSP enforcement moved into the proxy with nonce-based script-src. The stale static CSP came out of the Next config. JSON 404s replaced HTML 404s for API paths. Public-facing copy got the remaining preview language stripped out.
The Arkus website: stop being an AI demo, start being a funnel
The website did the most direct course correction this week.
The public Arkus AI APIs are gone. The chat endpoint, the live-token endpoint, the TTS endpoint, the assessment endpoint, the chatbot component, the voice agent component, the assessment services, the Gemini client dependencies. All removed. Arkus AI is an access-gated platform page now, not a public live demo.
In their place: a hidden Apply funnel layer, backed by Notion. Attribution captured. Lead scoring rules applied. Trial requests routed to manual review. Newsletter signups routed to the Insider waitlist. Dedupe in place so partial submissions do not create duplicate Notion records.
The week did open with one piece of remediation work. The live-token endpoint had been returning the wrong kind of credential for a browser surface. It now mints short-lived, constrained tokens through GoogleGenAI.authTokens.create, scoped to the live model and voice config.
The website's job is to qualify and route. The AI product's job is to run for the people who belong inside it. Those are different jobs, and they should not share the same surface.
What is next
More direct-source grounding on patent claim retrieval. Continued boundary work on Index workers, including the queue and retry behavior. Continued portal work on VentureIP as real members move through it. A cleaner runbook for model quota and billing behavior so routing decisions stop being made under pressure.
Back next week. Same repos, still in flight.