Tags on this page use a fixed taxonomy so filters stay stable over time:
Cloud, CLI, Platforms, OSS, Dashboard, Registry, Billing,
Community, and Breaking change.npm and Bun CLI installs now follow the verified Desktop-managed runtime when
it is the same version or newer, so a Desktop update also updates normal
terminal commands. The packaged runtime also discovers its bundled DuckDB
resources without relying on Desktop-only environment variables.Source checkpoint: Desktop now shows native update availability in the sidebar
footer, with a download icon that reveals “Update” on hover or keyboard focus,
download progress, retry feedback, and the existing restart confirmation.Source checkpoint: macOS update restarts now enter window shutdown before
Electron closes windows, preventing close-to-tray behavior from blocking a
downloaded update. This fix is not included in the published 0.4.10 release.Source checkpoint: agents can load an exact local skill revision, activate a
selected subset or collection for one task and harness, inspect unfinished
activations, and clean up receipt-owned links. Changes preview by default.
Repeated activation is idempotent; overlapping tasks and replaced targets are
guarded. Interrupted tasks remain inspectable rather than expiring on a timer.
These commands have not yet been published in a packaged release.Source checkpoint:
selftune skills search adds offline BM25 discovery across
installed skills and inactive Library packages, with bounded JSON results and
collection membership. Search does not activate or download skills; availability
in packaged releases is tracked separately in GitHub releases.Skill Set license drafts now show preview and apply errors inside the draft dialog.
Long license diffs scroll inside a bounded review pane, keeping the dialog and
its review buttons within the window.
Switching Sets or revisions resets the sharing dialog so license drafts use the
displayed Set’s skills rather than a previous selection.
Imported Sets can draft licenses for their pinned Library packages without an
installed copy. Applying reviewed terms creates a new Set revision and preserves
the original package; changed revisions require a fresh review.Desktop runtime validation now accounts for code-signing changes to the bundled
DuckDB native libraries. Signed-bundle verification and exact checks on copied
runtime files remain required. Release availability is tracked separately in
GitHub releases; a source checkpoint is not a published download.npm artifacts omit development snapshots, Storybook examples, and README animations.Release verification can generate the npm candidate’s SBOM again. The bundled
dashboard packages declare one
lucide-react range, the shared UI package pins its
diff and tree rendering libraries to the versions it is tested with, and the SBOM is
generated and validated with a locked CycloneDX toolchain before publishing to npm,
so a failed candidate cannot leave npm ahead of the GitHub release again.
Runtime SQL migrations and their journal remain included and are checked after installation.Team workspaces can now configure signed HTTPS webhooks for authoritative
assignment creation and reported installation-status changes. Endpoint
management is owner/admin scoped, secrets are shown once and stored as a
SHA-256-derived signing key, payloads are privacy-safe and bounded, and the
durable outbox exposes stable delivery IDs, retry attempts, and failures.
This was verified locally without sending to a real endpoint and is not a
production-delivery claim.Team owners and admins can now create explicitly scoped, revocable service
credentials for a versioned Skill Set release API. The raw token is returned
only at creation and only its hash is stored. Publish intent/finalize,
promote/deprecate, and bounded release status reuse the same Team entitlement,
release authority, and audit trail as the product.The new noninteractive
selftune team CLI drives that authoritative release
and assignment lifecycle without reviving the retired Registry state model.
Mutations require --yes, automation can request stable JSON and exit codes,
and service or linked-device credentials are accepted only through an
environment variable or standard input—not command arguments or logs.A reusable, pinned GitHub workflow can now validate an exact reviewed Skill
Set envelope locally and emit deterministic revision and envelope hashes.
Promotion runs only behind the skill-set-promotion environment and submits
those exact hashes through the scoped service API; credentials stay in secrets
and environment variables. Tests made no external call or production release.Team owners and admins can now create scoped service credentials using
plain-language presets, review safe credential metadata, and revoke access.
Cloud shows the raw token exactly once with a save-now warning and copy action;
lost tokens cannot be recovered and must be replaced.Desktop now offers managed Cloud’s supported one-time copy link without
exposing reusable-link, email, complete-backup, Shared Packs, or workspace
rollout controls that belong to a different backend. Unknown Cloud API paths
return a JSON 404 instead of the browser application, and retired hosted
product routes render as unavailable rather than opening the workspace.Cloud now also supports workspace switching and actionable people, device,
and one-time-share management. Active Team seat limits count joined members
and pending invitations; role and resource controls are target-specific.
Lapsed plans block additions and promotions but preserve authorized cleanup,
and security activity uses a safe actor label instead of exposing account IDs.Team publishers can now preview and explicitly publish an exact Skill Set
release from Desktop. Cloud verifies the canonical object, keeps immutable
release history, and provides read-only list and detail previews with plain-
language contents, checks, tools, licenses, and provenance. Publish, share,
assign, download, and install remain separate actions.Team owners and admins can now assign an immutable release to one member’s
linked computer with Manual as the safe default policy. The
recipient reviews the exact local plan before installing and can undo through
the receipt-owned rollback path. Cloud shows desired versus Current, Failed,
Rolled back, or Unknown using bounded receipts; Unknown never counts as a
successful install, and local paths, prompts, sessions, and rollback contents
stay on the recipient’s computer.Team rollout status now reports how many assigned computers have returned a
bounded installation receipt, the resulting percentage, and grouped failure
codes. Computers without a receipt remain Unknown, including during partial
rollouts or network loss. Retrying the same receipt is idempotent, while
conflicting replay data, superseded assignments, and revoked devices remain
blocked.Managed Team contributions are now explicit. Desktop startup and scheduled
sync never package or upload local edits in the background; a teammate must
review and deliberately submit a candidate before it can reach Cloud.A recipient can now preview changes against the exact assigned release and
confirm a private package upload. Cloud derives the file manifest, review
diff, and readiness from the uploaded bytes; owners and admins can then
approve or reject the immutable proposal. Approval creates a new immutable
release but never assigns or installs it automatically. Confirmed offline
submissions retry from a private local outbox without creating duplicates.Desktop can also run an on-demand local collision check before rollout.
Trusted routing overlap and observed runtime regressions can block readiness;
text similarity is advisory and never treated as collision proof. Only the
bounded readiness summary is returned, while raw queries and sessions remain
local.Cloud now shows named Skill Set owners and stewards alongside healthy, stale,
or abandoned status. Readiness reasons and broad usage freshness are written
in plain language; raw prompts, sessions, and evaluation details remain local.
Owners and admins can manage stewardship by teammate name.Desktop publish review now shows the deterministic package lock and impact.
Explicit missing dependencies, cycles, incompatibilities, conflicts, omitted
components, revision drift, or a changed reviewed lock block publishing before
package bytes leave the device.Team owners and admins can explicitly promote or deprecate an immutable Skill
Set release. Cloud keeps an append-only lifecycle history and security audit
event; changing lifecycle status never changes release bytes, assignments, or
recipient installations.Assigned Skill Sets now recompute local revalidation state when their exact
source, dependency lock, harness, or policy fingerprints are observed. Cloud
receives only Ready, Needs review, or Could not test for the exact
assigned member and linked computer; model names, reason details, and raw test
evidence remain local. Offline summaries retry without changing their reviewed
lifecycle payload.SelfTune now models AGENTS.md, CLAUDE.md, Cursor project rules, and Copilot
repository instructions as explicit adjacent instruction types. Each declares
its own import, validation, placement, activation, conflict, update, and Undo
behavior. Unsupported harness semantics fail closed, and these files are never
re-labeled as Agent Skills or uploaded to Cloud.Authoritative Team assignments now support Manual, Notify, and
explicitly opted-in Automatic update policies. Existing Ask before
updating assignments map to Manual. Automatic updates fail closed and return
to recipient review when local conflicts, stale bases, missing local policy
evidence, incomplete readiness, an unpromoted or deprecated release, or a
dependency-lock mismatch is detected.Cloud sign-in now pairs the account form with a responsive preview of skill
sets, project-scoped context, and secure sharing. Email sign-in, account
claiming, verification recovery, and password reset remain in one focused
journey, with clearer field labels and direct links to the Terms of Use and
Privacy Policy.
selftune mcp serve now exposes the local skill catalog through the draft
Skills-over-MCP discovery methods and read-only search_skills and
load_skill tools. Entries are pinned to immutable revision URIs and carry
SHA-256 resource manifests. Loading instructions does not install skills,
execute code, or grant tool permissions.The public CLI, Desktop app, embedded skill, and release artifacts now share
the published
0.3.3 baseline. Pending release notes are owned by the
workspace-visible Desktop release train and resolve to the coupled 0.4.0
release for the local-first skill management relaunch. The public export now
carries its browser-safe API contract and standalone test dependencies, so a
clean public checkout can verify the release without private-monorepo state.
The temporary Next.js infrastructure shell now adapts explicitly into the
shared journey modules while the greenfield Cloud Host owns product UI.
Server-side review exports now load only the portable review contract.
Documentation now starts with the complete skill-management journey—discover,
clean up, package, project-scope, share, and maintain—and uses the relaunch
ink, cream, and lavender palette instead of the retired cyan theme. Human
guides now focus on outcomes and decisions, while a separate agent section
provides operational contracts, approval boundaries, stop conditions, and
expected results. Oversized guides were split into focused setup, scoping,
updating, and sharing journeys; current branded Library and Skill Set visuals
replace the old cyan screenshots; and unreleased one-off execution choices
are no longer presented as available product actions.The hosted companion is now a deliberately small Convex boundary for
identity, subscriptions, linked devices, privacy-safe inventory, explicit
package sharing, consented contributor-signal aggregates, update signals, and
security history. The OSS self-host now implements the same Desktop state,
manifest, and contributor-signal journeys for a customer-operated trust
boundary without Stripe or PragSys’s cross-customer SaaS control plane.
Desktop remains the
product and keeps skill contents, agent history, evaluation evidence, and
improvement execution local. The former Fly, Neon, hosted telemetry,
cloud-runner, sandbox, and remote-improvement source paths have been retired.Existing hosted users can keep their workspace without carrying old provider
sessions or password hashes into the new system. SelfTune now offers GitHub
sign-in plus an email-verified account-claim and password-reset journey. The
official Convex Resend component queues delivery durably, prevents duplicate
sends, and automatically removes finalized delivery records after seven days.
Workspace administration now follows that same retained boundary. Desktop’s
Team view shows who has access, each person’s linked devices, installation
coverage, harness and scope metadata, aggregate last-used activity, update
drift, and clear adoption or maintenance recommendations. Raw prompts,
sessions, evaluations, and skill files remain local. Settings owns member
administration: owners can invite people by email and assign roles, and
pending invitations are accepted automatically after the invited person signs
in. Resend delivers branded invitations from SelfTune’s verified notification
domain; membership changes remain scoped to the linked workspace and are
written to its security history.Codex rollout and wrapper ingestion now removes recognized internal context,
environment, and orchestration blocks before recording user queries or skill
attribution. Wrapper-only messages no longer become user evidence, while real
user text following a recognized wrapper remains available for analysis.
The Mac app now detects whether the SelfTune agent skill is installed before
configuring history, hooks, and operating mode. When it is missing, onboarding
provides a bounded copyable prompt for the user’s AI agent and an explicit
manual install command. Installation remains user-directed and does not enable
hooks, import history, or connect cloud services automatically.Packaged Desktop builds now retain and verify the complete local DuckDB
runtime before launch, preventing incomplete app bundles from reaching the
onboarding flow.
The homepage, quickstart, navigation, evaluation concepts, creation guide,
improvement guide, and CLI reference now share one product story: observe
local skill behavior, package realistic evals, compare the same cases against
a baseline, and keep only evidence-backed changes. Cloud remains documented
as an optional sharing and distribution boundary rather than the default
improvement runtime.
selftune eval run now executes each packaged case with the current skill and
a no-skill or previous-skill baseline in separate working directories. Each
iteration records outputs, timing and token measurements, assertion evidence,
blind comparisons, benchmark deltas, and a human-feedback artifact.Skill packages can now keep routing cases, realistic output-quality prompts,
expected results, and mechanical assertions under
evals/. SelfTune prefers
these versioned package artifacts over local cached copies while retaining
compatibility mirrors for existing workflows.Desktop and Cloud now compose the same Skills, Skill Sets, Plugins, sharing,
and collaboration journeys through small explicit modules. The former
whole-product host adapter and its compatibility projections have been
removed, while authentication, navigation, and transport remain owned by
each host. This simplifies product changes without changing the existing
local data boundary or Cloud authorization model.
A new Plugins page detects existing Claude and Codex installations on the
current machine, groups the same plugin across both hosts, and shows versions,
enabled state, source, ownership, and version drift. It includes search and
host/ownership filters and distinguishes plugins installed by SelfTune from
plugins installed elsewhere.Management stays inside each host’s supported CLI boundary. Claude plugins
can be updated, enabled, disabled, or removed when their scope permits it;
Codex plugins can be removed. Removal requires confirmation, and Claude data
is preserved for a later reinstall. Organization-managed Claude plugins remain
visible but defer controls to the host policy that installed them.
DashboardOSSPlatforms
Desktop installs a reviewed Skill Set in Claude and Codex with one confirmation
Desktop now detects the official Claude and Codex CLIs and can install one
exact sealed Skill Set revision into either or both hosts. The confirmation
shows host availability, current installed versions, the projected plugin
version and revision, skill count, user scope, and activation instructions.SelfTune materializes a revision-addressed local marketplace, then delegates
registration and installation to each host’s official CLI. It never edits
Claude or Codex registry files directly. A stale revision fails closed, and
the initial native installer carries skills only; MCP capabilities remain a
separate review boundary. Cloud and self-hosted pages can hand a Pack to
Desktop but cannot change local plugin state from the browser. The Codex
action is local and does not claim publication into ChatGPT web or a
workspace directory.
Desktop now shows a labeled Delete action on saved Skill Sets. Deletion
requires confirmation, removes the Set from the local library, and leaves
existing project installations unchanged. Immutable revision history and
cached skill packages remain available for audit integrity. Sync & Backup
respects the local deletion instead of restoring the Set from remote history.
Cloud and Desktop Skill Sets can now export their exact pinned revision as a
Claude plugin, an OpenAI plugin, an Agent Plugins 1.0 package, or one
all-formats archive. Agent Plugins output uses the official versioned root
manifest and keeps available SelfTune revision and BOM provenance in its
extension namespace.Exports re-verify every pinned component and produce deterministic ZIPs. The
download flow explicitly distinguishes an archive from a direct-install URL.
A new fail-closed source-composer contract supports exact pins from Skills.sh,
GitHub, local folders, and archives without silently omitting unresolved
components.MCP remains separate from Skill Sets. An optional capability attachment must
be bound to the exact set revision and explicitly reviewed for commands,
network access, credentials, and persistent plugin data before it can be
projected into a plugin.Skill Sets can also be shared as branded Pack URLs. Cloud uses
cloud.selftune.dev/p/<opaque-id> and self-hosted installations use their
configured public origin. Desktop’s Add from URL flow previews skills,
license terms, expiry, and the immutable revision before import, then
re-verifies the downloaded envelope and every embedded package. Pack links
are expiring and revocable; unlisted links are reusable and private links
are consumed by the first successful download.The branded browser preview now shows the actual Set, components, licenses,
access mode, expiry, and revision, with a one-click handoff into signed
SelfTune Desktop builds. Desktop opens the review automatically but still
requires explicit import confirmation. A new Shared Packs workspace lets
owners copy active links, see claimed/expired/revoked state, and revoke access
without using a terminal. Cloud and one-container self-host expose the same
management behavior.SelfTune can now prepare bounded evidence cohorts from real historical agent
sessions and run managed before-and-after skill evaluations without copying
holdout evidence into the candidate. The evaluation path records exact skill
revisions, paired trial results, uncertainty, infrastructure failures, and
review-only recommendations before any deployment decision.Codex source sync now projects historical rollouts into both the canonical
SQLite store and DuckDB analytics while the Desktop dashboard is running.
The dashboard releases its DuckDB writer connection between OTLP requests,
preventing an idle daemon from blocking CLI ingestion and repair.Historical evaluation can now select neutral, exact local tasks, generate
several bounded skill-body candidates, calibrate them, and choose from a
Pareto frontier before running paired current-versus-candidate replays. The
managed Codex replay is read-only, records bounded process evidence, and
returns a concrete before-and-after result without applying the candidate.
Local task-quality drafts remain local replay artifacts and cannot be
submitted as correlated-error Cloud evidence.
The public site now presents SelfTune as the operating loop for teams that
need their skills to stay synchronized: capture changes, review evidence,
adopt the right revision, and keep managed workspaces aligned. The revised
pages make the team-sync workflow, product proof, and manager controls easier
to understand before a visitor starts a trial.
When a skill or Skill Set cannot be shared because license evidence is
missing, SelfTune now offers a Draft missing license action. Authors enter
the copyright holder and licensed organization, then review the exact
SKILL.md and LICENSE changes in a Pierre file tree and split diff before
approving them.The initial template grants internal use, modification, and private team
distribution while prohibiting external redistribution. Previewing never
writes files, cancellation leaves the package unchanged, and approval is
rejected if the skill changed after review. The draft is presented as a
drafting aid rather than legal advice. Public plugin export now reuses the
same validation boundary and keeps generated Claude plugin archives scoped to
the explicitly authorized skill package.Desktop startup now recognizes the narrow DuckDB index-replay failures that
can make a local observability WAL unreadable. SelfTune preserves the
checkpoint and WAL in a private, timestamped recovery backup, restarts from
the last healthy analytical checkpoint, and reports the backup location. It
does not apply this recovery to unrelated database errors.Revised analytical batches now checkpoint their indexed replacement
transaction immediately to prevent the same replay hazard. The Desktop reset
action also includes the DuckDB checkpoint and WAL in its recoverable state
backup while continuing to preserve user configuration. Recovery diagnostics
now distinguish the preserved checkpoint/WAL pair from unrelated reset data.
SelfTune now detects stable local edits to workspace-managed skills, packages
and deduplicates them, and sends them to the team’s review queue automatically.
No terminal command or agent prompt is required. Each candidate is pinned to
the contributor’s installed base revision, with a package hash, file manifest,
source identity, change summary, and exact added, changed, or removed file list
for review. Automatic teammate suggestions are explicitly marked Not
evaluated with No efficacy evidence attached, and approval publishes the
reviewed bytes without claiming measured improvement.
selftune registry suggest remains available as an advanced recovery command
when automatic delivery reports a failure.Registry suggestions now carry a durable outbox receipt through delivery and
installation, so operators can distinguish an evaluated recommendation from
a delivery retry or a client that has not acknowledged the rollout yet.Workspace admins can adopt or reject candidates in the shared Team
Collaboration dashboard. Adoption uses compare-and-swap publication, so a
candidate cannot replace a registry head that changed while it was waiting.
Adopted candidates can be rolled back to their exact base revision.Per-skill rollout policy can be manual, notification-only, or automatic. New
workspace skills default to automatic, so an owner’s approval is the last
human step before the next scheduled sync. Clients first hash the installed
directory and refuse to overwrite local edits. Managed devices send bounded
update, conflict, failure, and rollback receipts. Suggestions contain
managed skill package files, manifests, hashes, and a change summary; raw
prompts, transcripts, and usage content are not included. Aggregate usage
signals remain a separate opt-in channel.Local orchestration now scans a bounded set of fresh and historical
interaction signals after source sync, redacts their evidence, and prepares
review-only correction hypotheses for existing skills. A durable checkpoint
advances the historical scan without starving new interactions. Existing
correction records such as “why didn’t you use this skill?” are recovered
through the same bounded, idempotent path instead of being stranded in the
legacy signal table.New Write/Edit hooks capture the exact whole-package revision immediately
before and after a successful
SKILL.md change. The durable artifact contains
only revision hashes, timestamps, and a path digest; it stores neither skill
contents nor filesystem paths. Codex, Claude Code, OpenCode, and Pi feed this
evidence into the same correction-study discovery path.The Skill Library now includes an in-context correction review surface for
durable candidates. It shows the observed failure, correction intent,
bounded proposed diff, evidence level, verifier provenance, blind result,
regressions, and limitations. Accept, reject, and defer receipts are
append-only and idempotent; they never apply a skill automatically, and an
edited candidate must be captured and evaluated as new evidence.Ambiguous signals and evidence without exact skill revisions remain deferred.
Managed replay and blind comparison may promote a reproducible correction or
expose a review-ready candidate, but preparation alone never claims an
improvement and no proactive candidate edits a skill automatically.Skill Portfolio cleanup now recommends a never-invoked package for archive
review after both the inactivity window and subsequent-session threshold
pass. The evidence window begins at package modification time when no trusted
invocation exists.Protected, routing-problem, and consolidation candidates remain excluded.
Every recommendation still requires explicit review and moves the package to
reversible quarantine rather than deleting it.Bulk archive review now submits the selected packages through one local
operation, invalidates Library reports once, and refreshes once. If the
archive succeeds but that refresh is unavailable, the review closes and
reports the completed archive separately from the refresh failure.Desktop diagnostics also stop writing to a development console after its
output pipe closes, while file logging remains active, so a logging
EIO or
EPIPE cannot recursively bring down the local runtime.The maintained Cloud app now includes workspace billing and seat management,
account settings, Team administration, GitHub App connections, skill detail,
improvement runs, and proposal review. These journeys use the shared typed
Cloud API boundary rather than retired Cloud pages.Team invitations reserve purchased seats, preserve workspace roles when a
recipient signs in, and keep member changes restricted to workspace
administrators. GitHub installation callbacks use short-lived signed state
bound to the originating workspace and return only to the maintained GitHub
settings page.
The shared Cloud and Desktop dashboard now use calm cream surfaces, charcoal
controls, and semantic evidence colors for information, review, successful
improvements, and failures. Product components use shared design tokens
rather than fixed palette values, so both hosts stay visually aligned while
preserving the former dark direction as a temporary comparison theme.
Each Skill Library row now includes a 30-day trigger sparkline built from
trusted observations. Hovering the line shows the selected day’s count, the
30-day total, and the lifetime total. Skills with no recorded triggers remain
visible as a flat zero baseline without storing synthetic zero-value history.
Claude Code, Codex, OpenCode, Pi, and OpenClaw now own durable-source
adapters behind one shared registry. Sync, watch, and orchestration resolve
those adapters through the same boundary while preserving each harness’s
existing checkpoints and retry behavior. Adapter execution now stays inside
Effect through orchestration, with Promise conversion limited to CLI and
Desktop boundaries.Desktop setup and connection detection use a lightweight registry for all
shipped harnesses, including Cline, without loading session parsers. Parser
code loads only when a live source sync runs. Cline remains hook-only because
it does not yet expose a durable session source.
Supported Codex execution patterns can now select a deterministic, bounded
cohort of failed and successful runs without copying full transcripts into
SelfTune. Selected excerpts are allowlisted, redacted, byte-limited, and
separated into calibration and blind holdout evidence.The cohort enters the existing body-evolution lifecycle for the exact
installed skill revision and produces a minimal, provenance-rich body
mutation for human review. Blind holdouts never enter proposal generation,
stale revisions and insufficient contrast remain explicit diagnostics, and
generation never creates a new skill or applies the mutation automatically.
The shared Desktop review surface shows the evidence summary, exact diff,
evaluation state, uncertainty, provenance, and accept, edit, reject, or defer
actions.
Claude Code transcripts, Codex rollouts, OpenCode sessions, and Pi sessions
now project bounded metadata through the same Effect Schema and
LocalTraceImporter boundary. Each accepted source revision writes
replay-safe spans, five automatic metrics (duration, input tokens, output
tokens, errors, and tool calls), and explicit links to canonical skill
invocations in DuckDB. SQLite retains canonical product state and successful
import checkpoints. One DuckDB connection serves a complete sync run, and a
source marker advances only after both stores acknowledge the revision.
A forward migration removes the deprecated, rebuildable SQLite trace
prototype tables; durable harness sources and DuckDB remain authoritative.Large local histories now stay bounded during import. Codex streams rollout
JSONL instead of copying entire files into memory, reports progress during
long backfills, and safely skips isolated oversized records while preserving
surrounding metadata. OpenCode reads its current message and part schema
in bounded chunks and excludes embedded diffs, file bodies, images, and tool
output from projection. Pi now recognizes its real lowercase read events
and global skill registry, while Claude retains promptless subagent execution
as observability spans without inventing prompt or skill facts.selftune ingest claude|codex|opencode|pi, normal sync, watch, and evolve
refreshes now use that same source-adapter path. Replays do not inflate
results, revised source files replace their prior analytical batch, and an
analytical failure remains retryable. SelfTune uses actionable-turn timing
when the source proves it and otherwise keeps one bounded session span rather
than inventing finer timing.Correlated trace signals continue into the existing Desktop Skill Sets view
and can produce minimum-sample, explicitly non-causal execution patterns.
OpenClaw remains canonical-only and Cline remains hook-only. This local trace
path does not change billing, cloud uploads, or the V2 telemetry contract.Desktop now computes its heavier Library, Portfolio, Insights, and Skill
Intelligence views in isolated workers, serves the last successful snapshot
while they refresh, and refreshes each view only when its own source data
changes. Live dashboard updates replace redundant polling while preserving a
slower correctness refresh for Skill Intelligence. Interrupted telemetry
uploads are also recovered automatically, and historical staging data is
re-uploaded through an explicit delivery cursor before retention removes it,
so local storage can shrink without treating an enqueue as delivery proof.
CLIOSSDesktopDashboard
Desktop turns processed skill history into reversible cleanup recommendations
Desktop setup now processes selected session history before completing and
sends users to a Cleanup ready checkpoint when trustworthy inactivity evidence
identifies archive candidates. The review opens a candidate-only Library view
with eligible skills preselected, shows the evidence before confirmation, and
moves approved packages into local selftune quarantine without deleting them.
Archived packages remain available from a dedicated restore view, while
unobserved and under-observed skills are never recommended for cleanup.
The Library also detects duplicate and outdated installations across projects,
proposes a source-current canonical revision, and shows every affected hash
and path before approval. Approved consolidations archive the original copies,
replace project copies with links to SelfTune’s immutable Library package, and
retain a durable rollback receipt. Consolidation never permanently deletes a
package, and a changed package makes the prepared review stale instead of
applying an outdated plan. Its browser-safe package entry keeps the Desktop
Library available without pulling Node-only hashing code into Vite. Existing
users now see duplicate-install recommendations in the Overview cleanup
checkpoint without repeating onboarding. A bulk review preselects only
source-confirmed revisions, summarizes the archive-and-link impact, holds
ambiguous revisions for individual comparison, and reports each result
independently so one failure does not stop the remaining safe cleanups.
The same workflow is now available through
selftune skills consolidate for
single-skill and source-confirmed bulk cleanup, with JSON dry-runs, explicit
--yes approval, per-skill durable receipts, and idempotent
consolidation-rollback commands for automation and agents.
The Skills Library now presents those recommendations as a dedicated review
control with counts and clear archive or consolidation context, preselecting
applicable archive candidates for bulk action without making the normal
Library table feel like a recommendation queue. Consolidation results now
link directly to the Desktop decision history, label the receipt action as
Undo consolidation, preview the links and archived copies that will be
restored, and let a just-completed bulk cleanup be undone from its exact
per-skill receipts.Desktop and Cloud now expose Share on each Skill Set. SelfTune materializes
the exact pinned set and its included skills as one portable package, verifies
license evidence for every component, and creates either a reusable link, a
first-recipient link, or a private email invitation. The recipient combobox
lists workspace members, accepts another email address, and can share with the
entire workspace. Hosted distribution requires a SelfTune Cloud workspace;
self-hosted remotes continue to support backup and restore. Recipient
telemetry remains off unless separately enabled. Development builds can
preview the Cloud conversion state without disconnecting an account. The
conversion action opens a feature-led preview of portable sharing, backup,
installation, and telemetry controls before offering to connect the device
to SelfTune Cloud. It leads with using the same pinned set across a person’s
own machines and sandboxes, then explains private sharing without repository
access and the limits of controlling readable files after installation. The
same feature-led dialog now appears when an unconnected Desktop user tries to
back up or privately share an individual skill, instead of hiding those
actions or sending the user to Settings without context.
Creating, updating, or deriving a Skill Set now requests Sync & Backup in the
background. Local saves still complete without waiting for Cloud, and the
existing scheduled sync retries changes that could not upload immediately.
The CLI and Desktop can now preview and configure multiple Skill Sets for one
existing project as a single conflict-checked operation. Desktop also offers
a guided React TypeScript project setup that applies selected sets after an
explicit preview and confirmation.
The Skills Library table now supports selecting individual rows or every
visible filtered row. A contextual toolbar can back up eligible selections to
Cloud, archive eligible skills after one confirmation, or clear the current
selection while preserving per-skill review flows for updates and removal.
The Desktop skill-details modal now shows one row per project root instead of
repeating a row for every agent-specific installation folder. Connected
agents appear as icons with a count, and local package records nested under a
known project are folded into that project instead of appearing as Unknown.
Cloud can now explicitly materialize an already-authorized Skill Set export
before download. The action reuses the exact source authorization and active
actor-bound lease, safely recovers expired attempts only while authorization
remains valid, and keeps ordinary export downloads read-only. Work that began
before authorization expiry may finish under its still-active exact lease;
completed sealed exports remain downloadable after expiry while their rights
and distribution stay live.
CloudRegistryCommunityBreaking change
Contributor signal routing now rejects unknown legacy destinations
Public and authenticated contributor signals, including community bundles,
now resolve legacy creator destinations only through a frozen server-owned
allowlist. Known previously accepted signal and bundle pairs continue during
the compatibility period; unknown, ambiguous, cross-organization, or sunset
destinations are rejected without deleting existing signal data. A route that
sunsets during ingest also rolls the payload write back atomically.
New Registry entries now default contributor feedback off across API,
GitHub, and Cloud Improve publication. An explicit publisher choice is kept
across later automatic version syncs; publishers can explicitly enable or
disable it on a later version. Existing Registry entries retain their stored
preference rather than being silently rewritten.
Connect Cloud account now opens
cloud.selftune.dev in the system browser,
returns to the pending Desktop connection after sign-in, and asks the user to
confirm the short-lived code before approving. Desktop stays local throughout
the flow and receives its own scoped credential after approval; it does not
import the browser’s Cloud session.Desktop now backs up linked Skill Sets and their pinned skill revisions when
installed packages use canonical symlink locations, without mistaking every
locally installed skill for a SelfTune release. Older learned-state records
remain portable, and a failed first backup stays visible in Settings with its
full diagnostic instead of disappearing with the connection toast.
A local skill can now be sent directly to the private Cloud Library from its
detail panel—without publishing, creating a draft, or wrapping it in a Skill
Set. After syncing another connected Desktop, choose a target agent and
install the verified whole-folder revision into its conventional global skill
directory. Existing unmanaged folders remain protected from overwrite.
The Cloud workspace now leads with two customer-facing destinations:
Skills and Skill Sets. The sidebar uses the shared shadcn workspace
selector, and the workspace home no longer exposes authentication, transport,
or shared-code implementation details.Hosted Skill Sets can be created, edited, downloaded, and removed directly in
the browser. Desktop-only filesystem actions and unavailable intelligence
views stay hidden until the active host supports them, instead of appearing as
warnings or empty tabs.
The greenfield cloud host now serves its shared Desktop-shaped experience and
direct managed account connection through Neon’s official browser client.
The account screens now use the same Base Nova shadcn tokens and component
layer as the shared Desktop-shaped product instead of loading a competing
global style reset.
Hosted Library and Skill Set operations pass the resulting bearer token
through an isolated Cloudflare Worker and are validated before reaching
current cloud services. This remains an internal preview and does not replace
the live customer domain or the optional, local-first Desktop workflow.
selftune desktop remains the primary experience and now consumes canonical
Library and Skill Set routes from a shared product package. A new greenfield
cloud host composes those same routes with browser authentication and
workspace context. Its first typed cloud routes support Library reads and
Skill Set creation, editing, export, and removal. The previous cloud-v2 and
Next.js interfaces are frozen for new product development, and the desktop
app remains fully usable without a cloud account.
SelfTune’s Effect workspaces now use the latest Effect tsgo patch with
TypeScript 7 across editor, local checks, pre-commit hooks, and CI. Contributor
installs also include pinned Effect source checkouts for fast, version-matched
implementation lookup by developers and coding agents.
The Skill Library now keeps update, merge, removal, and refresh actions visibly
pending while they run, then replaces completed review state with a clear
success or error message. Insights refreshes animate while scanning, and long
suggestion titles no longer displace their status labels.
Skill Sets included in Sync & Backup are now Workspace-owned and become
available to every workspace member after sync. Desktop labels Personal and
Workspace sets in Projects, adds a Settings → Workspace screen for member
roles and policy management, and supports email-first invitations that are
claimed when the recipient signs in.Workspace admins can leave a set Allowed, require a reviewed Apply action,
Block it, or mark it Required. Desktop re-checks the policy before changing a
project and retains the last verified rule for temporary offline use. Viewer
access remains read-only; applying Workspace Skill Sets requires Member access
or higher. Private recipient sharing remains available for immutable copies
sent outside the workspace.
selftune init and the desktop setup wizard now run the same convergent
setup engine: re-running is always safe, and each step reports whether it
was already satisfied, newly applied, or failed. Desktop remains fully local
and usable without an account; choosing Connect Cloud account opens a
short-lived browser approval where people can sign in or create an account.
Cloud API keys move out of
config.json into the operating system’s credential store (existing
installs migrate automatically), and approval selects account-backed
Sync & Backup and attempts the first backup without asking anyone to copy a
token. Each installation rotates only its own device credential, so linking
or reconnecting one computer does not sign out other devices. Raw transcripts
remain local, and self-hosted Remote Library setups keep working unchanged.
The Desktop sidebar now shows whether SelfTune Cloud is connected: choosing
an unlinked Cloud profile starts the same approval flow, while choosing a
linked profile opens Sync & Backup inside Desktop. Opening the hosted Cloud
dashboard is a separate, explicitly external action. After browser approval,
the native Desktop window returns to the foreground.Improvement candidates and proposal reviews now use the same unified diff
renderer as the rest of the dashboard, keeping line numbers, change counts,
expansion behavior, and visual treatment consistent across review surfaces.
The Skill Set editor now groups installed revisions under one canonical skill
result instead of showing duplicate search matches. When multiple active
revisions genuinely differ, an explicit version selector keeps that choice
available without cluttering the search list. Background Library refreshes
also preserve unfinished names, descriptions, connections, and selections
while the editor remains open.
The Skill Sets workspace now focuses on saved sets, suggestions, and outcomes.
Detected projects moved into the creation modal, where From Library and
From Project share one flow and selecting a project prefills its folder,
name, and connections. When a skill has different installed copies, the
chooser now leads with readable scope and connection labels while keeping
paths and content hashes as secondary technical context.
selftune watch now parses, validates, evaluates, reports, and exits through
one Effect-owned command lifecycle. Database access is scoped, synchronization
and rollback are explicit services, legacy flags retain their behavior, and
focused parity tests protect the existing CLI contract while making the watch
runtime independently testable.The local and desktop dashboards now separate My Sets, Suggestions,
Projects, and Outcomes into focused views. Saved sets open first,
project capture no longer displaces the library, and suggestions use a compact
master-detail layout that keeps skill roles and validation evidence available
without turning the page into a long stack of cards.
Run
selftune sets capture inside a project to detect its active supported
harnesses, deduplicate installed skills, infer a readable name, and pin the
result as an immutable Skill Set. The local dashboard now lists observed
projects with a direct Capture as Skill Set action. Identical recaptures
are no-ops, while same-name sets with different contents remain protected from
accidental replacement.Claude Code hook events now forward to the running selftune daemon instead of
booting a fresh process, loading modules, and opening the database on every
agent turn. Guard decisions and activation suggestions are relayed unchanged,
telemetry events are queued in per-session order, and hooks automatically fall
back to fully local execution whenever the daemon is unavailable — no settings
changes required. New installs also launch hooks with bun instead of node,
cutting per-event overhead to roughly twenty milliseconds; existing installs
pick this up by re-running initialization.
Dashboard reports (portfolio, skill intelligence, insights, library) are now
computed in a background worker and served from persisted artifacts, so pages
load in milliseconds — including the first request after a restart — instead
of recomputing for multiple seconds on every view. The skill-intelligence
response shrank from megabytes to kilobytes by summarizing matched session
signals instead of embedding them in full, and delivered upload artifacts are
now pruned on a retention schedule so the local database no longer grows
without bound.
Codex rollout sync now detects appended session files and atomically replaces
their prior batch-derived prompts, skill invocations, and execution facts.
Stable canonical IDs prevent replays from inflating usage, failed database
writes remain pending for retry, and live hook or wrapper evidence is preserved.
Local, Desktop, Self-host, and Cloud dashboards again expose the full shared
workflow: searchable Skill Set selection, connected-install details, branded
merge controls, conflict-aware update review with receipts, evidence-backed
archival recovery, and host-appropriate status and update metadata.
SelfTune now validates a discovered set by checking whether its internal skill
relationships recur in newer, unseen sessions, instead of requiring every
member to appear at once. Reusable skills can participate in several sets with
a contextual role and membership score, and Projects shows source provenance
plus held-out relationship coverage for every recommendation.
selftune sync now shares one scoped local database across source ingestion,
repair, contribution staging, run history, and optional alpha upload. Interactive
runs stream progress immediately, while piped and explicit JSON runs remain
parseable even when a non-empty dry run emits per-session previews or an
OpenCode source reports diagnostics.CLI, local service, Desktop, and self-host builds now share one package for
config schemas and environment-derived paths. The new Effect boundary validates
reads, writes config atomically with owner-only file and directory permissions,
and remains bundled and smoke-tested in clean npm installs.
Cloud and direct GitHub installs now validate skill names and versions before
downloading or writing packages. Registry state is schema-checked on every
read, and sync accepts only exact
.claude/skills/<name> destinations.
Traversal names, unconfined paths, and malformed state stop before network or
filesystem mutation.Launchd plists and systemd units now publish through an owner-only, synchronized
temporary file and atomic rename. Reinstalling tightens permissive definitions,
replaces unsafe leaf links without modifying their targets, and never activates
a partially written supervisor definition.
Clean npm installs now include the XML runtime used by Windows service evidence.
Release smoke tests import both lazy service-program modules from the installed
artifact, catching missing bundled dependencies before publication.
Linux service installation continues to request user lingering best-effort so
SelfTune can start at boot without an interactive login. Stop and uninstall no
longer disable this user-global setting because other user services may depend
on it. Legacy SelfTune markers are ignored as non-authoritative metadata. Users
can still change lingering explicitly after considering every affected user
service.
Desktop development now refreshes harness presentation changes without
restarting Electron or reacting to unrelated runtime refactors. A refreshed
sidecar preserves the active dashboard route. The Skills Library also
canonicalizes duplicate repository origins, the HMR server loads the Geist
font correctly, and Cline uses its official product mark instead of an
approximation.
Local, Desktop, Self-host, and Cloud once again expose the actions and context
needed to manage a skill from the Library. Rows include update and details
controls, category filters, clickable repository or local-folder sources,
connection identity, and revision context. The details dialog restores source
previews, agent-assisted three-way merges, durable removal, and archived-skill
recovery. Cloud keeps its proposal-aware Review, Run, Fix, and Setup actions,
readiness ordering, blocked-state guidance, and compact inventory summary.
The dashboard now calls the reusable collection surface Skill Sets while
preserving
/projects as a compatibility route. Local, Desktop, and Self-host
once again show evidence-backed overlapping-set suggestions, held-out
validation, feedback calibration, review and dismissal actions, and measured
before-and-after outcomes. The shared host contract now requires an explicit
intelligence capability so future cross-host screen work cannot silently omit
these features.Doctor, status, last, telemetry, guided quickstart, local dashboard launch,
local daemon and supervised service management, recover, badge, and SQLite snapshot export now use typed Effect command
definitions and the Bun Effect runtime while preserving legacy argument
handling for command families that still forward raw flags.
Remaining legacy command families now share one typed routing registry and
load only their selected implementation module, reducing startup work while
keeping Effect-owned commands under a single explicit authority.
Doctor, status, last-session insight, and guided quickstart are now fully
owned by that Effect tree, including the no-command status default. Flags and
operands that their legacy fallbacks silently ignored now fail validation;
malformed tokens also cannot hide behind global help or version flags.
The complete
eval family now has one typed Effect command tree for generate,
unit-test, import, composability, and family-overlap. Each action keeps its
valid flags, --out compatibility alias, output, and exit behavior while
rejecting ambiguous boolean assignments, automatic boolean negations,
malformed numeric or choice values, extra operands, and junk hidden after
help. Runtime eval modules now receive typed inputs instead of reading
process.argv; composability preserves SQLite-first telemetry loading, JSONL
fallback, and machine-readable report output.
Alpha upload and relink now share typed Effect command and runtime programs;
manual uploads preserve readiness guidance, JSON summaries, dry-run behavior,
and failure exit codes, while relinking preserves device-code approval,
credential replacement, and machine-readable progress output.
Uninstall now belongs to the typed Effect command tree and separates its
ordered cleanup plan, injected live capabilities, and command-local lazy live
adapter. It preserves the existing strict flag grammar, JSON result, cleanup
order, and Promise-compatible command surface while rejecting malformed
arguments before help or destructive work can run.
Team Registry push, install, sync, status, rollback, history, and list now use
the same typed command tree and a lazily loaded Effect client. Local validation
still runs before authentication or network access, response payloads are
schema-checked, and filesystem and HTTP capabilities are provided only at the
CLI runtime boundary. Direct GitHub installs now reject links and special
files, validate repository paths and refs before reading content, and stage a
complete replacement so a failed install leaves the previous skill intact.
Registry installation state now uses locked, atomic transactions, preventing
concurrent installs and syncs from silently overwriting newer local state.
Recovery validates destructive combinations before opening SQLite, badge validates
inputs before loading evidence, and export validates the complete table selection
before opening SQLite or writing files. Latency-sensitive harness hooks remain
outside the heavier runtime. Dashboard harness icons now come exclusively from
package-owned descriptors, allowing unused duplicate static marks to be removed
without changing the displayed connections. Daemon startup now acquires and
releases its server, manifest, and runtime lock as one scoped lifecycle, while
invalid arguments are rejected before local authentication state is created.
Native service commands now cancel pending OS control processes on interruption,
verify the supervised executable identity, and only report authenticated runtime details.
Linux service status now reads systemd manager state and MainPID directly, while
stop and uninstall reconcile loaded units even when their unit files have diverged.
Authenticated daemon health also reports the runtime owner, supervision mode, and
executable identity so Windows can recover an orphaned service without trusting
a stale manifest or terminating an unrelated process.
Windows service lifecycle commands now scan the exact loopback listener after
Task Scheduler stops, re-authenticate the complete runtime identity before any
instance-bound shutdown request, use scan-only final verification, and retain
task artifacts whenever the port cannot be proven clean.
Windows Task Scheduler control now uses structured inventory, stable numeric
state codes, post-action verification, and validated absolute System32 tools
instead of localized command output or ambient executable lookup.
Windows service mutations also share one per-user SQLite ownership lock across
config directories; the operating system releases ownership after a crash or
forced process exit, so abandoned PID files cannot wedge later repairs.
service doctor now reports the fixed current-user Windows lock state without
creating local service state. service repair-lock accepts no path, PID,
config directory, or force flag: it repairs only an exact stale pre-SQLite
generation after acquiring the SQLite owner and atomically installs the
permanent compatibility fence.
Interrupted legacy cleanup is now resumed from a SID-bound durable journal,
and receipt-backed artifacts are quarantined and verified after their atomic
move before deletion. Replacements are preserved, mismatches are restored
without overwriting newer files, and full receipt generations are revalidated
across task mutations.
CLI release checks are now advisory: they follow the matching stable or beta
npm channel, cache a verified registry version, and print the matching Bun,
npm, or skill-install command without replacing the active process or its
supervised service. Signed Desktop releases retain their native updater as the
only in-place automatic update path.
The typed daemon command now validates installation-specific nonces for
supervised Windows wrappers while rejecting them on direct daemon launches;
nonce evidence stays out of runtime manifests and is available only through
the authenticated health response for receipt-backed ownership checks.Applying a remotely backed-up Skill Set now keeps its automatic download and
verification state visible in the shared Projects screen. Desktop preserves
machine-readable Sync & Backup failures, distinguishes authentication,
offline, unavailable revision, and integrity errors, and reports how many
missing skills were downloaded and verified before the project changed.
Local Store, Library, Skill Intelligence, and Source Management now ship as
explicit packages with focused Effect service boundaries and runtime adapters.
CLI metadata and dashboard contracts are organized by lifecycle area, while
preserving existing command help, dashboard APIs, local data, and packaged
desktop behavior.
Local suggestion mining now uses skill source identity, stops presenting two
packages from the same source as a finished set, and merges mutually supported
co-usage relationships into overlapping multi-skill recommendations. Reusable
core skills can participate in several sets, while sets of three or more must
meet the newer held-out session quality floor before SelfTune recommends them.
Skill removal and conflicting Skill Set replacement now use the same
restart-safe approval lifecycle as agent-assisted source merges. Reviews show
every affected path, archive or backup destination, and recovery behavior;
approvals revalidate fingerprints immediately before mutation, while stale,
expired, or declined decisions leave content untouched. Approved actions retain
auditable quarantine or rollback receipts.
Release validation can now run the same Library update journey against Local,
SelfTune Cloud, and Self-host servers, while packaged Desktop receives its own
native Electron journey. Each target records explicit pass, skip, or failure
results alongside screenshots, traces, logs, and recovery receipts.
CloudOSSDashboard
The dashboard can now switch explicitly between Local, Cloud, and Self-host servers
The shared dashboard now identifies its active server and can save, test,
rename, remove, and switch between SelfTune Cloud and custom Self-host
profiles. Desktop always keeps one protected “This Mac” profile, injects its
Local bearer credential from the native process, and clears host-scoped state
before navigation so data from the previous server cannot flash into view.
Local source merges and Cloud improve runs now render intent, evidence, candidate
diff, decision, validation, and apply outcome through the same typed adapters. A
versioned Run Package provides a concise agent-readable export while redacting
credentials and machine-specific paths and excluding authenticated artifact URLs.
Authenticated agents can request either the full package or its compact summary
directly from the improve-run API.
Local, Desktop, Self-host, and SelfTune Cloud now render Projects through the
same host-neutral screen. Local installations retain creation, editing,
project capture, searchable Library selection, installation previews,
conflict review, apply, export, and rollback. SelfTune Cloud stores reusable
Skill Sets as portable Remote Library snapshots, supports create, edit,
delete, and JSON export, and points project installation back to Local or
Self-host where filesystem access is available.
Contributors and CI can now run one machine-verifiable Library update journey
through the real Local dashboard and API. Each isolated run records pass, fail,
or structured skip metadata plus logs, screenshots, the scenario source, and a
browser trace, while verified attached worktrees are never restarted implicitly.
A prepared source merge now receives a stable approval ID and remains staged
across page or process restarts until it is explicitly approved or declined.
Approval rechecks the source, installed files, and staged candidate before any
mutation; stale or expired reviews remain auditable and direct users to prepare
a fresh candidate. Repeated decisions are idempotent and cannot apply twice.
The Skills Library now has one shared implementation for inventory, search,
filters, source links, responsive tables, skill details, update diffs, loading,
empty, and error states. Local and Desktop, Self-host, and SelfTune Cloud supply
their data and optional source-update, merge, removal, and hosted actions through
the typed host adapter. Unsupported and upgrade-only actions are shown deliberately
instead of silently disappearing or branching inside the shared screen.
Accepting, rejecting, snoozing, or editing an Insight—and then drafting,
evaluating, or releasing it—now refreshes the queue, Library, overview, and
proposal views through one semantic contract shared with live events. Failed
attempts preserve stable cached state, and retries emit only the successfully
persisted decision.
Connection detection, Settings labels and visuals, and agent-assisted source
merge eligibility now come from each harness package. Renderer payloads use an
explicit safe projection that excludes functions, credentials, environment
values, and local configuration paths; unsupported actions disappear naturally
when a package does not provide their implementation.
Creating, editing, deriving, exporting, planning, applying, and rolling back a
Project Skill Set now use the same semantic resource declarations as live
dashboard events. Project materialization refreshes Projects, Library, and
overview together, while failed or stale revisions preserve the last stable
cached state and continue to surface their actionable error.
Repository contributors can now start, inspect, open, and conservatively reap
a watched Local runtime with Vite HMR through internal
bun run dev scripts.
Each worktree owns an authenticated port block and manifest, preventing parallel
checkouts from attaching to or stopping one another. These are development
scripts and are not added to the customer-facing selftune CLI.Local, Cloud, and Self-host dashboards now contribute authentication, queries,
mutations, navigation, permissions, and feature availability through one typed
host adapter. Shared screens can behave consistently without branching on where
they run, while unavailable and upgrade-only features remain explicit.
Source previews, agent merge preparation, applied upstream updates, removals,
restores, and live server events now share one typed resource registry. An
applied change refreshes the Library, portfolio, overview, intelligence, and
related Project state together, while failed staged operations retain the
last stable view and partial removals still reconcile every derived screen.
Settings and the desktop tray now call the customer feature Sync & Backup and
clearly distinguish SelfTune Cloud from a self-hosted server. The destination
selector, connection state, sync actions, privacy note, success messages, and
errors use the same language: the Library remains local, while selected
immutable artifacts are copied to the configured destination. Connection,
preview, integrity, missing-object, and uninstall messages now name that
destination consistently. The Remote Library name remains unchanged for the
deployment-neutral protocol, API routes, configuration files, and engineering
interfaces.
Sync now discovers hosted Skill Set manifests that are not yet present on the
current device without replacing local sets. When Apply finds a missing pinned
revision, SelfTune downloads that exact immutable package from SelfTune Cloud
or the configured self-hosted server, verifies its object and package hashes,
stores it in the local Library, and only then links it into the project.
Conflicts, offline destinations, missing objects, and integrity failures stop
before project mutation. Fully local applies continue without a network call.
The same compiled
selftune binary now owns daemon startup, its authenticated
durable manifest, and launchd, systemd user, or Windows Task Scheduler
registration. The desktop app is a thin host for that runtime: it monitors
health, performs bounded recovery, exposes restart and diagnostics actions,
preserves a backup before resetting runtime state, and keeps the recovery
screen available when startup fails. Runtime manifests now distinguish the
CLI or desktop owner from desktop-child, OS-service, or unsupervised lifecycle
authority. The app preserves CLI foreground daemons and equal, newer, or
unversioned services; it upgrades only a provably older registered service,
while authenticated desktop children orphaned by an app crash are reclaimed.Signed releases still update through GitHub Releases. Packaged runtimes are
copied out of transient install media into a versioned stable path only after
their signed-source manifest verifies. The packaged-app release proof launches
the final Electron bundle, exercises the authenticated Settings dashboard and
preload bridge, verifies the stable runtime copy, and confirms graceful child
cleanup, including quit during an in-flight startup. Installed Unix runtimes
must remain executable as well as byte-identical to the signed source. CLI
and desktop versions advance together, and release tags must resolve to the
exact source commit before any artifact is published. Bot-authored Version
Packages pull requests explicitly dispatch their required checks when no
separate release token is configured, so protected branches do not stall the
release loop. Release intent now enters through a workspace-visible package
before the public CLI and desktop versions are stamped together, and CI runs
each Effect suite with its native test runner. Cross-platform desktop proofs
also use native temporary paths and a Linux-only sandbox waiver confined to
the unpacked smoke process. Windows verification now treats Unix permission
bits as a POSIX-only contract and exercises service definitions with
host-native paths. Remote Library URL normalization and email-like invocation
detection use linear scans so untrusted input cannot trigger regex
backtracking. Public desktop and self-host builds now pin Bun 1.3.14. The
Windows runtime is cross-compiled on Linux, passed to the Windows packaging
job as a same-run artifact, and then exercised on Windows, avoiding the
upstream Windows-host compiler crash without weakening the platform smoke.The public repository now uses explicit Executor-style package ownership:
apps/cli composes commands, apps/local owns the daemon and HTTP host,
harness integrations live in independent packages, and cross-harness sync
lives in an orchestration package. Existing npm and hook paths remain as
compatibility shims. Source sync is provided through a typed Effect service
and live Layer, so reusable runtime science no longer reaches into platform
adapters through hidden imports. Inside the local host, dashboard workflows
now run through a scoped Effect service with typed failures and deterministic
disposal. Authentication, live events, SPA transport, protocol validation,
and SQLite-backed read routes have independent owners, leaving the Bun server
as a small process composition root instead of a second application layer.
The local operational store now uses Drizzle over Bun SQLite, with a typed
schema, committed migrations, a data-preserving gate for existing installs,
and an embedded migration payload for the compiled desktop runtime. The
long-running host acquires and releases the shared database through an Effect
Layer, while existing prepared query paths remain compatible during the
repository conversion.
The release packer flattens shared Effect and Zod runtimes into the root
manifest and normalizes bundled workspace links before npm snapshots the
tree, keeping the artifact within its enforced size and file-count budgets.Electron supervision now follows the same ownership model internally. A
scoped Effect runtime serializes restarts, service toggles, resets, recovery,
update preparation, and shutdown instead of dropping a request while another
transition is active. Health probes and child-exit events are tied to their
connection generation, authenticated window replacement is staged before it
becomes active, native tray payloads are schema-decoded, and disposable IPC,
updater, tray, session, and monitor resources are finalized by their owners.
Packaged builds now compile internal workspace TypeScript into the Electron
main bundle instead of asking Node to execute source files from node_modules.The local dashboard now uses a quieter application shell centered on Library,
Projects, Insights, and Settings. A shared command palette makes skills and
destinations searchable from the sidebar, inventory and project controls use
denser native primitives, and the desktop bridge can open verified skill
directories directly in Finder without exposing arbitrary file paths.
Source updates now expose the recorded-base, local, and latest-upstream diffs.
When both sides changed, users can choose Claude Code, Codex, OpenCode, or Pi
plus an optional model override to prepare a true three-way merge. SelfTune
validates and stages the agent result for review, then rechecks the source,
local files, and candidate before a separate backed-up apply step.Dashboard pages now share one visual scaffold across Library, Projects,
Insights, Settings, and Status. Library prioritizes essential columns on
narrower windows, Insights groups secondary review actions, loading states
mirror their finished layouts, and an available dashboard build appears as
an Update action beside the sidebar version instead of a floating alert.The local Library now classifies installed skills into stable semantic
categories using package content and local observation evidence. Projects
suggests reusable Skill Sets when trusted sessions show a repeated ordered
workflow, strong unordered co-usage, or a recurring project toolkit. Every
suggestion separates discovery evidence from recurrence in a chronological
holdout, labels the result exploratory, supported, or validated, suppresses
stale patterns that do not recur, suppresses sets that already exist, and
opens the normal review dialog before creation. The same local,
LLM-free analysis is available through selftune sets suggest --json; raw
transcript text is never included in the report or uploaded.
Category corrections now retain the inferred label alongside the user’s
choice, and Projects records accepted, edited, and reasoned dismissal outcomes
against immutable evidence snapshots in the local Drizzle store. Temporary
dismissals can return only after the aggregate evidence changes, while
structural rejections suppress the same recommendation.
Calibration is versioned and remains inactive until at least 20 usable human
reviews include both positive and structural negative examples. Accepted sets
are measured against matched before-and-after project sessions across
completion quality, errors, trigger coverage, token use, and grading; small or
mixed samples remain inconclusive, and the product makes no causal claim.
Decision-history sync now restores redacted overrides, corrections, evidence
snapshots, reviews, and outcome aggregates across devices while raw
transcripts and session bodies remain local.On macOS, the desktop window now places the native traffic-light controls
directly beside the SelfTune wordmark instead of reserving a separate title
bar. The sidebar and main top edge remain draggable, while interactive
controls continue to receive clicks normally.Remote Library account tokens now use the macOS Keychain, Linux Secret
Service, or Windows Credential Manager when available. Existing plaintext
configuration migrates automatically to a credential reference; headless
environments keep an explicit owner-only file fallback.OSS SelfTune now includes a production one-container self-host for immutable
backups, multi-device access, and recipient-scoped private sharing. The
non-root image runs the canonical dashboard and cloud-compatible Remote
Library API with tenant-scoped SQLite and content-addressed objects in one
/data volume. Snapshot heads use transactional compare-and-swap, account
tokens are hashed at rest, raw transcripts never sync, and coordinated backup
instructions cover both SQLite and object content. Release automation builds
linux/amd64 and linux/arm64 images for GitHub Container Registry, smokes
the immutable per-platform digests under a non-root read-only container, and
moves version or latest references only after the exact candidate passes.
The self-hosted web Library and Skill Sets views read the authenticated admin
organization’s remote snapshot, with Skill Set mutations left to trusted
synced devices. Rollback receipts also recognize when a filesystem reuses an
old device and inode, so SelfTune preserves a replacement package instead of
treating it as the materialization originally created by a Skill Set apply.Remote Library now stores one canonical immutable revision for a skill even
when that revision appears in several locations or harnesses. Skill Set sync
includes every pinned revision in the same snapshot, so a restored Set can no
longer reference a package that was never uploaded. Evaluated release
authority remains attached to the canonical package when available, and old
released_skill snapshots remain restorable.Desktop Settings can grant a skill or Skill Set revision privately to an
existing SelfTune user. A Skill Set grant is expanded and validated on the
server, recipients must explicitly accept it, and import copies every
immutable object into the recipient organization before advancing its
snapshot. Senders can revoke outstanding grants, expiration is enforced, and
create, accept, import, and revoke actions enter the audit log. Device API
keys never act as share links.The supervised desktop service now performs Remote Library sync shortly
after startup and every four hours while it remains configured. Raw
transcripts still never enter the backup or sharing protocol.Installed macOS builds can now keep SelfTune’s authenticated local service
running independently of the Electron window and menu-bar process. An
owner-scoped LaunchAgent starts at login, restarts only after failures, keeps
its bearer token out of the plist, and is repaired when an app update changes
the bundled sidecar. SelfTune falls back to a managed child if supervision is
unavailable, and the menu bar can enable or disable the service explicitly.Signed desktop releases now use GitHub Release update manifests. The app
checks on launch and every four hours, downloads updates in the background,
shows progress in the menu bar, and asks for an explicit restart after the
update is staged. The release workflow publishes blockmaps and merges macOS
arm64 and x64 metadata so each Mac receives the correct build.The Library harness column now uses compact logo-only indicators with
accessible hover and focus labels. Pi receives a larger internal treatment
so its tall mark has the same apparent weight as the other harness logos.Library loading now uses cached source-update metadata for the first render
and refreshes GitHub state in the background. Repeated harness links share a
single package hash per scan, while portfolio combination analysis avoids
processing unrelated query logs when it only needs positive examples. The
native shell also uses a compact summary endpoint instead of downloading the
full local analytics history on every page.
The native app now organizes SelfTune around Library, Projects, Insights, and
Settings. Library reconciles duplicate installations, immutable revisions,
drafts, cached packages, and reversible archives across Claude Code, Codex,
OpenCode, OpenClaw, and Pi. Project Skill Sets can target all five harnesses,
preview every materialized path, block conflicts before mutation, prefer
links to a verified cache, and roll back only receipt-owned paths.Insights adds a local synthesis inbox for repeated successful coverage gaps
and stable ordered skill combinations. Candidate strength includes
independent-session support, project diversity, temporal recurrence, outcome
quality, marginal co-usage lift, sequence consistency, confidence, and
uncertainty. Supporting and held-out sessions are separated before draft
creation. Accepted candidates create drafts with workflow, provenance, and
positive, negative near-neighbor, boundary, and execution eval cases; they do
not publish, install, merge, or archive anything automatically. Explicit
evaluation binds package validation, replay, routing, no-skill and existing
skill baselines, held-out lift, and regression checks to the candidate,
evidence snapshot, and immutable draft hash. Release is blocked if that draft
changes or the gate does not recommend it.Library inventory rows now show the installed source, global or project
location, configured harnesses, trusted last invocation, package modification
time, and upstream update state. GitHub update badges are backed by install
lock metadata and concurrency-bounded repository tree checks; local and
unsupported sources are labeled as untracked or unavailable instead of being
guessed.Source checks now reuse a six-hour disk cache and authenticate with
GITHUB_TOKEN, GH_TOKEN, or the active GitHub CLI session before consuming
anonymous API quota. The Library detail panel exposes every concrete path,
scope, harness, source, revision, last use, and modification time. GitHub
updates are previewed against the recorded upstream tree, block on local
changes by default, and require an explicit replace action that retains a
complete backup and update receipt. Inventory discovery also vendors all 73
placement definitions from the upstream vercel-labs/skills registry. This
expands where SelfTune can find skills without claiming telemetry support for
agents that do not yet have a SelfTune adapter.Remote Library is now an optional deployment-neutral backup and multi-device
protocol for immutable packages, selected drafts, Skill Sets, metadata, and
decision history. Raw transcripts remain local and are not a supported sync
artifact. A sync preview lists exact artifacts and byte counts; selected draft
provenance uses pseudonymous session identifiers and review reasons are
redacted before the remote boundary. Versioned Skill Sets can also be derived
from an existing project and exported as portable checked-in manifests with
no device paths or credentials. The same authenticated API runs in SelfTune
Cloud or in the optional OSS self-host. The shipped self-host uses
tenant-scoped SQLite and immutable filesystem objects in one volume, with
bootstrap accounts, health checks, integrity diagnostics, and coordinated
backup and restore instructions.The OSS CLI now inventories installed skills even when they have no usage
records, distinguishes missing evidence from measured inactivity, and
recommends keeping, measuring, repairing, consolidating, or reviewing a
package for quarantine. Approved quarantine operations move complete packages
outside active registries, preserve a local receipt and package hash, and
return an exact restore command. SelfTune and system/admin-managed skills are
protected, and no audit result deletes a skill automatically.The local dashboard now exposes that complete installed inventory directly,
including packages with no SelfTune observations, with search, evidence-state
filters, explicit quarantine, and receipt-based restore. A new Electron host
packages the same dashboard and local API as a native desktop control plane:
it supervises a compiled Bun sidecar on authenticated loopback, shares
~/.selftune with the CLI, and keeps the bearer credential out of renderer
JavaScript.Desktop Settings now reports whether Claude Code, Codex, OpenCode, OpenClaw,
and Pi are detected and whether their SelfTune integration is actually
connected. It also configures the three local automation jobs with guarded
human-readable schedule presets, recommended defaults, and native launchd or
systemd activation. The desktop bundle includes its own task CLI, so
background jobs do not depend on a separate global SelfTune installation.First-run desktop onboarding now lets a human choose historical import
sources, select the detected harnesses where SelfTune should install live
hooks, and choose between observability, daily health recommendations, and
autonomous improvement. These are executable controls: source choices govern
every sync path, feature choices reconcile the native background jobs, and
hook choices preserve third-party entries while installing app-bundled native
runners. Observability and recommendations are enabled by default;
autonomous improvement is explicit opt-in.On macOS, SelfTune now remains available from the menu bar after its dashboard
window closes. The live menu summarizes system health, connected harnesses,
and attention items; it can sync sessions immediately, toggle each native
automation job without changing its configured frequency, open Settings,
status, and logs, and enable Launch at Login in installed builds. Action
failures surface as native notifications, while an explicit Quit action shuts
down the supervised local sidecar cleanly.Frequent telemetry sync now imports new harness sessions without rebuilding
the complete historical repair overlay every 30 minutes. The daily health
job retains the full repair checkpoint before scoring and recommendations.SelfTune now stores reusable project Skill Sets in a content-addressed local
Library. The CLI and native Projects screen can pin installed package
revisions, preview Codex and Claude Code project links, block the entire
operation on any destination conflict, apply the set idempotently, and roll
back only paths owned by the apply receipt. The desktop navigation is reduced
to Library, Projects, Insights, and Settings so inventory, distribution,
evidence, and configuration each have one predictable home.Positive-control qualification now accepts only a versioned operator canary
whose current skill, corrected candidate, and 50-case deterministic suite
match registered content hashes. The Worker and independent campaign
assessor reject missing or drifted control identity before treating results
as evidence, so an ordinary generated candidate cannot masquerade as proof
that the measurement instrument detects real improvements.
Cloud improve, discovery, qualification, and sandbox experiments now check the execution provider’s
key allowance and account credits before reserving a source or dispatching a
workflow. Unavailable or unverifiable capacity returns a retryable service
error without dispatching provider work; if database cleanup is interrupted,
a queued idempotency retry must repeat admission before it can dispatch.
Late provider failures remain excluded from scored evidence. Qualification protocols also seal an
operator-declared campaign reserve and both execution ceilings. Scheduling
and Worker start require that reserve on a finite, non-resetting execution
key; account-credit checks use a separate management key that is never
passed to the task-agent sandbox. Sandbox experiments use their own finite
execution key, while a global account-capacity lease still serializes all
provider consumers because both keys draw from the same credits. Campaign
reserves use exact cent arithmetic, reject an unexecutable final repeat, and
require prior repeats to pass in order. Duplicate experiment delivery is
atomically claimed, stale coordinator leases expire, and exhausted dispatch
recovery atomically fences workflow start before releasing both leases. A
durable cleanup state retries lease release after crashes or coordinator
failures. The zero-work infrastructure control stays runnable during
provider outages.
Large outcome suites now fail before entering the ordinary atomic improve
path. Operators can instead run a bounded calibration-only discovery that
persists candidate archives and diffs without declaring a winner or creating
a proposal, then stage the candidate as a detached snapshot for sealed
qualification. Qualification checkpoints every task attempt independently
and reserves explicit runtime headroom so long repeated trials can resume
without discarding earlier evidence.
Claude Code hooks now hash the complete installed skill package used by an
invocation, persist that version and invocation time through local staging and
cloud upload, and bind replayed tool events to the prompt active when they
occurred. Historical events without preserved package bytes remain explicitly
unversioned and are excluded from causal improvement claims.
Task-package runs now remove verifier scripts, oracle data, declared hidden
assets, prior answers, and stale rollout artifacts before the task agent
starts. The sealed archive is restored only after the agent exits, and newly
authored reviewed-trace packages no longer include the previous response.
Cloud operators can preregister immutable current-skill, candidate-skill, and
no-skill experiments with randomized arm order, repeated scored trials, and
sealed holdouts. Each task agent now hands only its source subtree to a fresh,
network-denied verifier sandbox, while the protocol seals the exact Worker
version, runtime manifest, model order, locale, and repeat count.
Qualification runs cannot create winners, proposals, readiness changes, or
apply attempts. Their full evidence is content-addressed in R2, and campaign
assessment verifies every stored byte against immutable Neon hashes before it
can pass.
Processed raw telemetry pushes remain available in Neon for seven days, then
move to R2 under deterministic keys with SHA-256 evidence. The cloud worker
verifies each object before clearing the duplicated hot payload columns, and
archived pushes remain replayable through the hosted API.
CloudDashboard
Cloud experiments now expose SelfTune-native agent trace observability in the dashboard
Completed cloud experiment runs now retain aggregate agent trace evidence,
including turn counts, tool-call counts, token totals, timing fields, and
final-response status for evaluation review.The dashboard also has an experiment trace view that shows the sandbox boot
snapshot, ordered lifecycle events, category summaries, and expandable raw
event payloads without relying on third-party tracing UI.
Cloud
Cloud sources can now draft and queue cited benchmark-factory arms from uploaded skill snapshots
The hosted API now exposes a source-scoped benchmark factory draft endpoint.
It reads the current skill snapshot, extracts cited contract requirements
from the skill body, proposes distinct scenario candidates, and returns
deterministic verifier proposals with explicit draft-only promotion blockers.The hosted API can also review completed sandbox arm evidence for the same
source-scoped benchmark draft, validating verifier proposals against
current-skill and no-skill outputs before emitting review-only promotion
drafts.Benchmark-factory drafts can now queue bounded current-skill and no-skill
sandbox experiment arms. Each queued arm carries scenario, condition, and
batch metadata so completed experiments can be reviewed without losing the
baseline identity.Completed queued arm experiments can now be reviewed by batch or experiment
ID, converting persisted sandbox traces back into arm evidence before running
verifier validation and discriminating-case filtering.
CloudDashboard
Cloud V2 adds the first Hono-backed Effect Schema contract endpoint for the Start removal refactor
The hosted API now exposes a Cloud V2 status contract endpoint backed by Hono
and Effect Schema. This is the first tracer bullet in the Cloud V2 refactor
away from TanStack Start server functions toward API-owned contracts and
backend Effect modules.The Cloud V2 overview page now has a Hono-backed readiness batch endpoint, so
the dashboard can replace per-source Start server-function fanout with one
API-owned request that returns null for individual missing readiness rows.Cloud V2 skills inventory and skill detail reads also now have hosted API
endpoints, giving the dashboard a Hono-owned read path before the Start BFF
layer is removed.Cloud V2 improvement list/detail reads, skill source-file reads, and skill
improve-run queueing now use Hono endpoints backed by Effect services, moving
more operator workflows off TanStack Start server functions.Cloud V2 bundle list/detail reads and GitHub bundle creation now run through
Hono endpoints as well, so bundle pages no longer need the Start BFF wrapper.Cloud V2 GitHub import source creation and imported-skill auto-setup now call
hosted Hono endpoints, removing another Start server-function hop from the
skill import flow.Cloud V2 overview loading and readiness advancement now use hosted Hono
endpoints backed by API-owned inventory and Effect overview services.Cloud V2 GitHub settings now load through a hosted Hono read model instead of
a dashboard Start server function.Cloud V2 uploaded skill imports now post directly to the Hono import endpoint,
replacing the upload-specific Start route for skill source creation.Cloud V2 GitHub installation binding now has a hosted Hono endpoint backed by
an Effect service, preparing the GitHub App callback flow for the API-owned
backend boundary.Cloud V2 bootstrap now loads from a Hono endpoint, removing the app-local
Start/session database bootstrap path from the dashboard shell.Cloud V2 now builds as a Vite SPA instead of a TanStack Start app. The
dashboard no longer ships Start server routes, app-local Neon database code,
or BFF adapters; browser routes call the hosted Hono API boundary directly.Cloud V2 Hono contracts now use concrete Effect Schema DTOs for migrated
run, skill, eval, improvement, and readiness payloads, and the SPA decodes
key read responses before handing them to route components. Improvement
workflow failures now map from tagged Effect service errors instead of route
string matching.The Effect architecture guardrail now covers Cloud V2 Hono routes, service
contracts, and browser boundaries, preventing broad Effect imports, runtime
Effect usage in the SPA, and unknown response DTOs in migrated contracts.Cloud V2 dashboard API calls now send Neon Auth JWTs to the Hono backend for
local verification, avoiding per-request remote session validation. The
overview readiness endpoint also reads readiness rows in one batch, cutting
the 26-source readiness request from multi-second fanout to sub-second API
latency.Cloud V2 Hono and bundle paths now avoid redundant type assertions and
ambiguous fallback defaults, tightening the Effect-backed API boundary
without changing user-facing behavior.Cloud V2 workflow endpoints now share one Effect result-to-HTTP response
helper, reducing repeated route-level try/catch handling across eval,
experiment promotion, GitHub import, and bundle creation actions.Cloud V2 eval routes and run-detail DTOs now share typed request and JSON
helpers, reducing repeated route validation while preserving the existing
Hono response shapes.Cloud V2 experiment promotion now records explicit calibration/holdout eval
split metadata for reviewed task-package cases. Reviewed traces default to
calibration, while the API can intentionally promote a case as holdout for
blind regression measurement.
CloudDashboard
SelfTune Cloud now includes a public skill validator with safe archive handling and autofix downloads
The cloud dashboard now exposes a public
/validate
page for uploading Agent Skills packages without signing in. The page returns
spec validation, best-practice lint findings, package inspection details, and
an autofix download when deterministic corrections are available.Public validation and autofix uploads run through the shared skill validation
library with bounded archive processing, generated temporary paths, typed
upload errors, cleanup after each request, and basic per-IP throttling.The validator can also export an uploaded skill as a Claude Code plugin zip,
adding the required .claude-plugin/plugin.json
manifest and wrapping the skill under
skills/<skill-name>/ for plugin installs.CloudDashboard
Cloud V2 local smoke now supports dev auth and verifies the improve review loop without production login
Local Cloud V2 development can now enable dashboard dev auth with
DEV_AUTH=1 and
NEXT_PUBLIC_DEV_AUTH=1. The dev user is
seeded with an active alpha enrollment and a pro Dev Org so local smoke tests
can exercise the same upload, improve, proposal review, and apply surfaces
that production gates expose.The cloud dashboard smoke now has a local no-auth mode for dev-auth sessions,
waits for hydrated Cloud V2 surfaces before asserting, and accepts current
proposal handoff copy for applied improvements.CloudDashboardRegistryCLICommunity
Cloud private bundle creation now publishes from GitHub, registry updates are failure-safe, and telemetry projection can be replayed
The V2 bundle publish flow now starts from a connected GitHub installation,
repository, and skill path, then runs the private-bundle sync before the
bundle is shown as installable.Initial sync is required for new private GitHub bundles. If GitHub sync,
archive generation, manifest validation, or publishing fails, Cloud returns a
visible error instead of creating a success-shaped empty registry entry.Bundle lists and detail pages now distinguish installable releases from
bundles that still need a published version, so operators can see whether a
bundle has a current version before handing install commands to users.The CLI registry installer and sync command now verify downloaded archive
hashes before extraction, stage updates outside the live skill directory, and
swap the full directory only after extraction succeeds. Files removed from a
published version are removed locally on the next install or sync, while
failed updates leave the previous installed copy intact.Registry push and rollback persistence now compensates across DB and R2
boundaries: uploaded archives are deleted if DB writes fail, new versions are
not marked current until the switch succeeds, and rollback failures restore
the previous current-version pointer.V2 telemetry ingestion now fails visibly when raw payload archival or
canonical projection fails, leaving the push in a retryable failed state
instead of returning success with empty or missing artifacts. Contributor
signal, community bundle, and creator analytics reads are scoped to the
authenticated organization so guessed creator IDs cannot cross org
boundaries.Hosted Cloud Improve Worker persistence now follows the same retry-safe
shape as the API path for run candidates and search frontier state. Workflow
retries can fill or update candidate rows after a terminal run write, retain
useful specialists, shadow weaker winners, supersede weaker active frontier
entries, and prune stale retained frontier records.Proposal review routes now accept only approval or rejection. Applied state
is reserved for the explicit Cloud Improve apply endpoint, which preserves
the required human approval gate and records the actual draft or GitHub PR
apply attempt.
The hosted API now exposes the same explicit proposal apply endpoint and
cloud-run proposal filter as the dashboard API, so production smoke tests can
approve and apply Cloud Improve winners without direct database access.Operator alpha routes now include a projection reconciliation action for V2
telemetry. It revalidates archived canonical push payloads, replays the
dashboard projector with per-push source arrays, and marks failed projections
as retryable instead of hiding rebuild failures.The cloud-v2 improvement review page now uses the explicit Cloud Improve
apply service after approval and shows recorded draft or GitHub pull request
apply attempts, including failed attempts that need operator attention.The cloud-v2 run list, run detail, and candidate diff reads now load from
hosted Hono API endpoints with Effect Schema request contracts, removing more
dashboard-only server function reads from the operator path.Cloud-v2 improvement review, apply, and publish actions now call explicit
hosted Hono endpoints backed by Effect action services and stable typed error
responses, so the review page no longer depends on Start server functions for
proposal mutations.Cloud-v2 eval suite list/detail/save/smoke flows and experiment detail/promote
flows now use hosted Hono endpoints with schema-backed contracts, removing the
eval and experiment Start server-function wrappers from the dashboard app.Cloud-v2 import and bundle creation flows now share the same hosted GitHub
source-selection endpoints for installations, repositories, and skill path
discovery, reducing duplicated selection logic between the two surfaces.The Cloud-v2 browser API client now centralizes base URL handling,
credentials, JSON request bodies, 404-to-null reads, and stable error message
extraction for the migrated Hono endpoints.Cloud Improve run detail now shows a live progress strip and refreshes active
runs automatically, while terminal runs show the final state without polling.
Persisted run errors are also shown inline on the run detail page so failed
provider, storage, or projection states are visible to operators.
Run candidates now link to their proposal review handoff and show the latest
draft or GitHub apply attempt state when an apply has been tried.
Applied draft proposals can now be published from the improvement review page
as the next installable registry version, with a Cloud Improve source ref
linking the release back to the reviewed proposal.Bundle release history now shows installs per version, making adoption visible
at the same point operators inspect source refs and release summaries.
Bundle detail now separates current-version installs from stale installs so
sync adoption is visible before operators optimize the next release.Release checks now include
smoke:cloud-product-lifecycle,
a hermetic lifecycle smoke that chains Agent Skills package validation,
SkillsBench-style no-skill eval contract checks, distribution, CLI
install/sync staging, telemetry replay, Worker result persistence, and
approved proposal apply guardrails into one JSON summary.Bundle detail now exposes installable artifact provenance for the current
version, including source type, source ref, repo, archive size, and content
hash, with the same source/archive fields visible in release history. Release
history also shows aggregate pass rate, eval count, and session count per
version so adoption can be read alongside outcomes.CLICloudCommunity
Creator contribution setup now bundles a portable feedback helper for no-CLI downstream signals
selftune creator-contributions enable now writes
a generated selftune-feedback.mjs helper and
selftune.feedback.json manifest next to the
existing selftune.contribute.json config.Downstream agents can use that helper to submit one privacy-safe,
creator-directed signal after first-run consent, without requiring the full
selftune CLI or a contributor API key. Cloud accepts those helper submissions
at POST /api/v1/public/signals and stores them
under the creator’s organization.The V2 Cloud registry now applies the same helper automatically when a skill
bundle is published or synced from GitHub, with a publish-time checkbox for
teams that need to ship the raw skill archive without contributor feedback
artifacts.CloudBreaking change
Cloud eval cases move out of the legacy casesJson blob into the cloud_eval_cases table; the eval-suite smoke endpoint now returns `cases` instead of `casesJson`
The cloud_eval_suites.casesJson blob has been retired. Eval cases now live
in a dedicated cloud_eval_cases table that supports per-case identity,
lineage, lifecycle, and indexable metadata. The reader cutover and dual-
write that landed across slices 91-94 are now collapsed: writers go only
to the new table, the EVAL_CASES_TABLE_AUTHORITATIVE feature flag is gone,
and the column has been dropped.The wire-shape change to be aware of: the
POST /v1/eval-suites/:id/smoke
endpoint now returns suite.cases (the resolved case array) instead of
the legacy suite.casesJson field. API clients that unpacked
casesJson should switch to cases.CloudDashboard
Cloud improve runs now persist explicit search policy controls for candidate budget and mutation surfaces
Cloud improve run creation now accepts and persists search-policy controls,
including candidate budget and mutation surfaces, so the control plane and
Cloudflare runner execute the same optimization contract the product asked
for.This lets trigger-only runs stay focused on trigger-relevant surfaces while
outcome-backed task-package suites can intentionally open up body and
structure optimization.
CloudDashboardRegistry
Cloud moves private bundle publishing out of GitHub settings into a dedicated registry flow and trims settings back to connection management
The GitHub settings page no longer carries both installation management and
the full private bundle publish composer in the same screen.GitHub settings is now scoped back to installation binding, org-wide
write-back defaults, and connected repository sources, while the actual
private bundle flow now lives at
/registry/new.That new registry flow still uses a compact step-through wizard, but it now
sits under the information architecture the user expects: registry publishing
under Registry, GitHub access management under Settings.CloudDashboardRegistry
Cloud GitHub installs now use a signed start/callback flow so workspace binding survives the GitHub round-trip
The GitHub settings page no longer sends operators straight to a raw GitHub
install URL with no workspace context.SelfTune Cloud now starts the flow through a signed install-state redirect,
receives the GitHub setup callback on a dedicated callback route, and sends
the browser back to the originating workspace with the
installation_id intact.That fixes the broken state where the GitHub App could already be installed
on an account, but the workspace still had no bound installation row, which
left repository selection permanently empty.CloudDashboardBilling
Cloud billing success now reconciles completed Stripe checkout sessions directly instead of waiting indefinitely for webhook-only sync
The billing success screen now uses the returned Stripe Checkout
session_id to reconcile the completed
subscription directly into the current workspace instead of relying only on a
webhook-driven database update.That means local development and delayed webhook cases no longer leave the
workspace stuck on the indefinite “Setting up your subscription…” spinner even
when Stripe already shows the checkout as paid and trialing.The success page also now falls back to a bounded retry loop with an explicit
retry state instead of spinning forever after the sync window is exhausted.CloudDashboardRegistryBilling
Cloud GitHub bundle setup now renders a real upgrade state for free workspaces and uses a clearer staged publish flow
The GitHub settings page now checks workspace billing first, so free
workspaces see a direct upgrade state instead of a raw
402 API error when
GitHub publishing is not available on the current plan.The private bundle setup flow was also tightened into a clearer staged
experience that separates bundle destination, GitHub source selection, and
release behavior, with a compact publish summary alongside the form.That makes the page usable as an operator workflow instead of exposing the plan
gate as a transport error and burying the actual publish decisions in one large
settings form.CloudPlatforms
Cloud API background workers are now explicit so Neon compute can scale down when the product is idle
SelfTune Cloud no longer starts in-process API schedulers just because a
DATABASE_URL is present.Cron workers now require an explicit ENABLE_BACKGROUND_WORKERS=1 worker
deployment, the rollback-only legacy cloud-improve dispatcher stays off unless
CLOUD_IMPROVE_RUNTIME_MODE=legacy is set intentionally, and the Cloudflare
runtime recovery sweep now runs every 15 minutes instead of every 5 minutes.That reduces accidental always-on Neon compute usage while keeping the
background repair paths available when they are explicitly needed.CloudDashboardRegistry
Cloud org workspaces now switch in-app and private registry bundles can be bootstrapped from the selected workspace
SelfTune Cloud now lets operators switch org workspaces directly in the app
instead of implicitly binding the dashboard to the first org membership on
the account.That selected workspace now carries through the dashboard shell, registry
surfaces, and GitHub-connected private bundle bootstrap path, so partner-style
operators can move between client workspaces without dropping into a separate
browser session.The sidebar header was also tightened into a compact workspace selector with
inline search and workspace creation, while private registry detail pages now
double as install and onboarding surfaces for org-scoped bespoke bundles.
CloudDashboard
Cloud uploads now stay on the skill workspace after the first quick suite is prepared
Uploading a new cloud skill still prepares the first quick eval suite
automatically, but SelfTune now keeps you on the canonical skill workspace
instead of auto-starting the first run and pushing you into the run screen.That keeps the first improve run explicit, surfaces the setup and validation
state in one place, and makes the live run page a secondary diagnostics surface
instead of the default first impression.
CloudDashboard
Cloud run review now defaults to one compact summary and moves deeper evidence behind diagnostics
Completed cloud improve runs now lead with a single proposal-review pointer
instead of stacking release-gate, trust, timeline, candidate, and artifact
surfaces above the fold on the run page.The lower-level evidence still exists, but it now lives behind a Diagnostics
disclosure so the completed run screen reads as a lighter handoff instead of
an operator console.
The cloud source page now shows prepared validation fixes in a dedicated
review surface directly below the hero instead of burying the full
SKILL.md diff inside the setup card.That review surface renders the exact unified SKILL.md diff plus the
directory-change list in one place, so the apply decision happens next to
the real package changes rather than next to generic warning copy.The cloud source workflow now lets you apply the prepared deterministic
validation autofix directly from the default validation preview card.That means SelfTune no longer stops at “here are the warnings and the diff”
when the bounded next write is already known. It can create a refreshed
source snapshot for you and immediately re-run validation on that package.
CloudDashboard
Cloud validation now prepares package fixes with a visible SKILL.md diff and directory-change preview
When validation fails on a cloud source, SelfTune now prepares the
deterministic package autofix preview from the current snapshot archive and
shows it directly in the source workflow.That means the default source path now surfaces the proposed
SKILL.md diff
and directory changes instead of only listing validation warnings and making
the operator infer the next step manually.CloudDashboard
Cloud runtime details now separate active-loop coordination from lower-level source internals
The source page’s runtime lane now splits into
Run coordination versus
Source internals, so active-loop guidance stops sharing one mixed panel
with the lower-level source diagnostics.That keeps the source workflow closer to one guided hosted loop even after
you open the deeper technical surfaces.The source page’s benchmark authoring lane now splits into
Draft review
versus Bundle workflow, so task-package remediation follows a clearer
scaffold, materialize, and smoke-check sequence.That keeps the authoring lane closer to a guided workflow instead of one
broad technical tool surface once a runnable draft is in progress.CloudDashboard
Cloud source technical details now default to benchmark authoring only when a task draft is active
The source page’s technical details now default to
Benchmark authoring
when a persisted task-check draft exists, and to Runtime state otherwise.That keeps benchmark remediation work in front of you when SelfTune already
has a bounded authoring draft in progress, without making the runtime and
coordinator state fight for space in the same default technical view.CloudDashboard
Cloud source advanced workflow details now split trust rationale from technical authoring state
The source page’s advanced workflow surface now splits into
Why and trust
versus Technical details, so benchmark pressure, trust rationale, and
reviewed suggestion queues no longer share one undifferentiated panel with
coordinator and authoring state.That keeps the source workflow closer to one guided hosted loop even after
you open the deeper details.CloudDashboard
Cloud source pages now keep run pressure and uncovered-case queues behind the advanced workflow details boundary
The source page now keeps latest-run pressure, uncovered-case suggestions,
and reviewed suggestion history behind the existing advanced-details
boundary instead of showing that whole queue in the default setup flow.That makes the main source experience read more like one guided hosted loop,
while still keeping the deeper benchmark and trust evidence one click away
when you actually need it.
Cloud source detail now returns a suite-aware improve-run gate for the
selected eval suite, and source-detail refreshes preserve that suite id
across analysis reruns, validation polling, and latest-run refreshes.That removes another page-local trust inference path from the source screen
and keeps the run-start gate aligned with the same backend contract the rest
of the hosted loop now uses.
CloudDashboard
Cloud benchmark repair drafts now follow task-package workflows automatically when the source already uses them
When a source already relies on task-package suites, SelfTune now seeds the
auto-prepared benchmark repair draft as a task-package scaffold instead of
always starting from a generic structured check.That keeps bounded benchmark remediation aligned with the existing
task-package workflow earlier in the loop, including the later materialize
and smoke-check path.
CloudDashboard
Cloud hosted-loop state now distinguishes benchmark repair work in progress from review-ready drafts
When SelfTune has only auto-seeded the benchmark repair draft, the hosted
loop now reads that as
system remediation still in progress. It only flips
to AI remediation once the bounded refinement pass is complete and the
draft is truly ready for review.The source page follows the same split, so the run-start gate and next-step
guidance now distinguish “SelfTune is still working” from “now it needs
your review.”CloudDashboard
Cloud benchmark repair remediation now runs in the background for recent active sources
The same bounded benchmark repair drafting and one-pass refinement flow used
on source detail can now also run from the scheduler for recent active
sources.That means SelfTune no longer needs a source page read before it can prepare
the next review-first benchmark repair checkpoint for a source that is
accumulating trust pressure.
CloudDashboard
Cloud sources now block new improve runs until auto-prepared benchmark repair drafts are reviewed
When SelfTune auto-prepares a bounded benchmark repair draft, the source UI
and the run-creation control plane now both hold the next improve run until
that draft is reviewed.This turns the benchmark repair checkpoint into a real hosted-loop gate
instead of leaving it as a visible suggestion beside an otherwise still-live
Start improve run action.When SelfTune auto-prepares a bounded benchmark repair draft, it now also
runs one refinement pass through the existing authoring path before handing
the draft back for review.That means the first repair proposal is less skeletal, while the release
gate still stays in review-first mode instead of silently saving new
benchmark cases into the canonical suite.
CloudDashboard
Cloud source workflow headers now treat auto-prepared benchmark drafts as AI remediation checkpoints
When SelfTune auto-prepares a bounded benchmark repair draft, the canonical
source
hostedLoopState now reflects that as an AI remediation checkpoint
instead of still reading like the source is simply ready for another run.That keeps the source workflow header, release gate, and benchmark repair
card aligned with the same autopilot step rather than splitting the story
across separate local UI hints.When a cloud source already has a saved suite but recent run pressure or
pending eval suggestions show one clear benchmark hole, SelfTune can now
auto-prepare one bounded task/output repair draft instead of only warning
about the trust problem.The draft is persisted through the existing eval authoring session model,
recorded in the audit trail, and surfaced in the source workflow hero so the
operator can review the suggested repair directly.
CloudDashboard
Cloud source remediation now runs in the background instead of waiting for a page view
The first safe cloud source self-heals no longer depend only on someone
opening source detail.SelfTune now runs bounded background remediation for recent active sources,
so ended Apply Observation windows, stale saved-suite smoke checks, default
judge calibration drift, and stale source coordinator locks can be repaired
before an operator even refreshes the page.
The cloud source setup hero now reads the backend
hostedLoopState
contract for its top-level loop summary and next-action copy instead of
relying only on page-local wording.The selected-suite trust gate still appears as a separate setup blocker
underneath that header, but the primary workflow frame now stays aligned
with the same stage/status model that already drives run and proposal pages.CloudDashboard
Cloud source pages now show when SelfTune already fixed deterministic blockers for you
Cloud source detail can now clear two deterministic blockers before the page
even asks for operator input: it settles ended Apply Observation windows,
reruns the latest eligible saved-suite smoke check against the current
snapshot when it is missing or stale, reruns the default judge calibration
benchmark when it is missing, stale, or context-drifted, and clears stale
source coordinator locks after the persisted Cloud Improve Run has already
ended.The setup hero now shows a
SelfTune auto-remediation card so operators can
see what the system already handled automatically instead of guessing why the
trust state changed.Hosted improve run detail and proposal detail now read the backend
hostedLoopState contract for their top-level loop summary and next action
copy instead of relying only on page-local wording.That means the workflow headers now stay aligned with the same stage/status
model the API emits, while older fixtures still fall back safely until every
surface is fully migrated.The shared
Automation and release gate card now shows the next ladder rung,
Trusted automation, instead of stopping at the current release boundary.Operators can now see whether a source is simply too early, not ready because
trust is holding the gate, or still on the assisted-default path because
higher autonomy is a later policy tier.CloudDashboard
Cloud source, run, and proposal pages now explain what SelfTune automates versus what still needs approval
Cloud source setup, improve-run detail, and proposal detail now share an
Automation and release gate card instead of expecting operators to infer
the product boundary from separate trust panels and button states.The card makes the current Assisted default mode explicit, says what
SelfTune already handles automatically, and says what still needs human
approval before anything ships.CloudDashboard
Cloud setup heroes now say when setup is complete but trust is still blocking the next run
Cloud source pages no longer show a flat
Ready to improve state when the
setup checklist is complete but trust history or another release gate is
still blocking the next run.The primary setup card now explains that setup is complete while the next run
is still blocked, and surfaces the blocker inline instead of burying it under
the disabled button.When a cloud improve run is blocked by trust, the setup hero now explains
what to do next instead of only showing a disabled button.The primary card surfaces direct actions like
Show advanced details and
Review eval suite, so operators can jump straight to the place where the
blocker can be resolved.CloudDashboard
Cloud source pages now keep the default setup loop ahead of advanced trust and authoring controls
Cloud source detail now keeps the workflow hero and suite editor at the top of
the page, and moves coordination, task-package authoring, structure
recommendations, and full trust panels behind one explicit advanced toggle.The default path now reads more like
review checks -> start improve instead
of forcing operators to scan the whole optimization console before taking the
first action.Cloud source setup, improve-run detail, and proposal detail now share the
same hosted-loop header before the deeper trust and evidence panels.The pages now reinforce the same
Connect -> Prepare -> Improve -> Review -> Apply -> Observe model so operators can see where they are in the loop and
what happens next without learning the underlying optimizer internals first.CloudDashboard
Cloud improve runs now join winning search provenance and trust evidence in one review card
Improve run detail now shows one
Optimization context card for the winning
candidate instead of forcing operators to mentally combine the search notes
and trust panels.The card combines winning search provenance, saved-suite smoke, eval-vs-
observed outcomes, judge calibration, and trust-history context in one
review surface.Proposal detail now shows one
Optimization context card that ties the
chosen search strategy to the suite trust state that should govern whether
the candidate is believable.The card combines search provenance, saved-suite smoke, eval-vs-observed
outcomes, judge calibration, and trust-history context so operators can
judge the optimizer and the measurement story together.Cloud source detail now treats a
Hold trust state as a real run-start
barrier, not just a summary badge.When the current saved suite or measurement trust is too degraded to trust
another run, the quick-start hero disables Start improve run and explains
which stale trust driver is blocking the action.Proposal detail now treats a
Hold trust recommendation as an actual apply
barrier instead of only a warning.When the surrounding saved-suite smoke, source trust, calibration, or
outcome history says the proposal should not be trusted yet, the apply
actions are disabled and the page explains that trust must improve first.Proposal detail now turns the surrounding trust context into a simple apply
recommendation instead of only listing caution bullets.Operators now see whether the page believes the proposal is
Ready,
Review carefully, or Hold, based on the suite smoke result, source trust,
calibration health, trust-history advisory, and rollback-review signals.Proposal detail now surfaces trust cautions directly in the apply panel when
the surrounding measurement state deserves extra scrutiny.Before applying a proposal, operators now see warnings for
watch or
stale source trust, failed saved-suite smoke checks, weak or drifting
judge-calibration results, and recent Apply Observation windows that are
still pending.Improve-run detail now reuses the same source trust context that proposal
review already shows.Hosted run pages now keep the saved-suite smoke result, judge-calibration
history, and trust-history ledger visible next to the run evidence so
operators can review search output and trust state in one place.
Cloud proposal detail now shows why a candidate was selected and why its
measurement should be trusted in the same review flow.Proposal pages now combine bounded search provenance with the matching source
trust context: eval-suite trust status, latest saved-suite smoke result,
judge-calibration benchmark history, and the source trust-history ledger.
Cloud source trust panels now include a chronological trust-history ledger
instead of relying only on summary buckets.The ledger combines saved judge-calibration runs, canonical saved-suite smoke
checks, suite-backed Cloud Improve Run outcomes, and Apply Observation
outcomes into one operator-readable timeline.
Cloud source trust panels and observed-skill cloud controls now surface the
latest saved
llm_judge calibration benchmark, the change versus the
previous comparable run, the weakest benchmark families, and recent benchmark
history.That gives operators a concrete review workflow for prompt/model calibration
drift instead of relying only on the raw calibration CLI output.Skill-level and source-level cloud trust panels now add coarse 30-day
outcome buckets across the last 90 days on top of the short weekly history.That gives operators a broader trust read before they queue another hosted
run, instead of relying only on the most recent few outcomes.
Skill-level and source-level cloud trust panels now group recent suite-backed
runs by each suite’s latest canonical saved check state.That makes it easier to see whether suites that currently look healthy are
actually lining up with helped outcomes after apply, or whether regressions
are clustering around failed or missing saved checks.
Skill-level and source-level cloud trust panels now include a compact
multi-window post-apply history derived from recent completed observation
buckets.That makes it easier to see whether the last few windows were improving,
regressing, mixed, or steady instead of relying on a single rolling badge.
Skill-level and source-level cloud trust panels now include a compact recent
outcome timeline based on the same post-apply observation summary that powers
the direction badge and latest outcome link.That keeps a short run of concrete helped, regressed, or inconclusive
proposal outcomes visible without drilling into proposal history.
Skill-level and source-level cloud trust panels now condense recent
post-apply outcomes into a compact direction signal:
Improving,
Regressing, Mixed, Steady, or Needs more signal.That makes it easier to tell whether trust is getting better or worse
without opening each proposal outcome one by one.The skill-level and source-level cloud trust panels now surface the latest
completed post-apply outcome with a direct proposal link.That means stale or watch-mode trust warnings now point at concrete outcome
evidence instead of leaving operators to hunt through proposal history by
hand.
Source trust summaries no longer wait for proposal-detail reads or the
batch observation scorer to reflect ended apply windows.When an observation window has already ended, the cloud trust summary now
scores it on read and promotes it into the completed outcome counts
immediately.
Cloud trust summaries no longer flatten everything into completed outcomes.
The selected source trust card and run preflight now show when recent applies
are still inside the post-apply observation window, so operators can tell
when the latest outcome counts are not final yet.
CloudDashboard
Observed-skill cloud controls now warn before queueing runs with stale or failing trust signals
The observed-skill
Cloud Improve panel now raises explicit run preflight
warnings when the selected source trust is already degraded or when the
selected suite’s last canonical task-package smoke check failed.That keeps risky benchmark state visible at the actual queue point instead of
only inside the deeper source detail screens.The
Cloud Improve panel on observed skill pages now shows the selected
cloud source’s trust summary, recent post-apply observation counts, and the
latest canonical task-package smoke result for the selected suite.That means operators can check benchmark freshness and recent real-world
outcomes before they queue another hosted improve run, instead of drilling
into the cloud source page first.Suggested trigger cases and recent run-pressure cards can now promote
directly into a review-only task-package draft instead of always starting on
the structured path first.That direct promotion path now follows the first persisted draft/refine step
with initial environment and verifier asset generation, so operators start from
source evidence plus concrete draft files instead of an empty task-package
scaffold.
Suggested trigger cases and recent run-pressure cases no longer create
task/output drafts purely in page state.Draft creation now goes through a review-first promotion route that builds seed
enrichment on the server and invokes the same Think or fallback refinement path
used by the persisted authoring session.
Cloud source pages now show the recent review-first authoring steps attached
to a persisted task/output draft instead of only the latest draft snapshot.The authoring session now keeps a bounded activity log for draft saves and
promotions, Think or fallback refinement, task-package asset generation, bundle
materialization, runnable smoke checks, and draft clears.
CloudDashboard
Cloud advanced run settings now show canonical task-package smoke freshness in suite selection
The advanced run drawer no longer hides whether a saved task-package suite
most recently passed or failed its canonical smoke check.This reuses the same summary-level saved-suite smoke signal in the advanced run
suite chooser, so operators can compare suite trust while configuring a run,
not just while editing the suite.
Operators can now compare saved task-package suites before opening one in the
editor.This adds summary-level canonical smoke state to hosted eval-suite list
responses and surfaces it directly in the cloud source page suite picker, so
saved-check freshness is visible during suite selection as well as after a
suite is opened.
CloudDashboard
Cloud source setup summaries now show the latest canonical task-package smoke result outside the editor
Operators no longer need to open the eval editor to see the last saved
task_package smoke result.This now surfaces the latest canonical smoke state directly in the cloud source
page setup summary, so the trust signal is visible in the broader read model as
well as in the editor.Running a saved canonical
task_package suite smoke check no longer yields a
result that disappears as soon as the request ends.This now writes the latest canonical smoke result back into the saved suite
metadata and surfaces it in the cloud source page editor, so operators can see
the last saved-suite check before deciding whether to rerun it.CloudDashboard
Cloud source pages can now run saved canonical task-package suites once against the current snapshot
Saved canonical
task_package suites no longer require a separate improve
run just to verify the saved case still works against the current source
snapshot.This adds a narrow saved-suite smoke action on cloud source pages and matching
eval-suite routes, so operators can run a saved canonical task-package case
once before they reuse it in hosted improve runs.CloudDashboard
Cloud task-package saves now retain draft provenance and latest smoke results in the canonical case payload
Runnable
task_package saves no longer drop the context that produced the
draft.This now writes a typed task_package_metadata block into matching canonical
cases, preserving the seed evidence, expected-outcome scaffold, optional notes,
and latest smoke result that came from the authoring session.CloudDashboard
Cloud task-package saves now require a fresh smoke result before a runnable case can be written into a canonical suite
Runnable
task_package drafts on cloud source pages can no longer be saved
into a canonical hosted eval suite if the latest smoke result is missing or
stale.This now blocks the save both on the source page and in the eval-suite
create/update routes, so runnable task-package cases must be freshly
smoke-checked before they become part of the authoritative suite.CloudDashboard
Cloud task-package drafts now keep the last smoke result visible and mark it stale when the scaffold, assets, or bundle change
Runnable task-package drafts on cloud source pages no longer silently lose
their last smoke-check result when the scaffold or generated bundle changes.This now keeps the latest smoke result visible, marks it stale with an explicit
reason, and tells the operator when the draft needs to be smoke-checked again
before save.
CloudDashboard
Cloud source pages can now smoke-check runnable task-package drafts before saving them into a canonical suite
Runnable task-package drafts on cloud source pages can now be executed once
against the current snapshot before they are saved into a canonical suite.This persists the latest pass/fail result back into the authoring session, so
operators can verify the materialized bundle and current snapshot still work
together before promoting the case.
CloudDashboard
Cloud task-package drafts now switch into an explicit runnable state once their review-only bundle is materialized
Materialized task-package drafts on cloud source pages no longer stay labeled
like scaffolds.This now promotes them into an explicit runnable draft state, updates the
promotion preview and bundle preview accordingly, and makes it clear that the
existing save flow will write a real canonical
task_package case.CloudDashboard
Cloud task-package drafts can now materialize review-only bundles into real R2-backed environment archives
Persisted task-package drafts on cloud source pages can now turn generated
bundle files into a real review-only environment archive.This uploads the bundle to R2, points the draft scaffold at the materialized
archive, and keeps the archive descriptor in the same authoring session until
the operator explicitly promotes the draft further.
CloudDashboard
Cloud task-package drafts now roll generated asset files into a review-only bundle preview with explicit archive-materialization readiness
Persisted task-package drafts on cloud source pages no longer stop at raw
environment/verifier asset text.This now rolls the generated files into a review-only bundle preview, marks
when the draft is ready for archive materialization, and keeps the whole flow
inside the persisted authoring session until an operator explicitly promotes it
further.
CloudDashboard
Cloud task-package drafts can now generate review-only environment and verifier asset drafts through the authoring agent
Persisted task-package drafts on cloud source pages can now generate a
review-only environment manifest draft and verifier script draft through the
authoring agent.This keeps the new task-package authoring flow behind the same persisted draft
session, surfaces whether the generated assets came from Think or the fallback
template path, and avoids writing anything canonical until the operator decides
the draft is ready.
CloudDashboard
Cloud source pages now support real task-package draft authoring, including editable scaffold fields and canonical task-package case saves
Persisted task/output drafts on cloud source pages can now move past a
placeholder preview into real task-package scaffold authoring.This adds editable instruction, environment, verifier, oracle, skill mount, and
resource-hint fields to the review-only draft flow, lets operators save that
scaffold back into the persisted authoring session, and makes the existing save
path emit a real
task_package case when that is the chosen promotion target.CloudDashboard
Cloud source pages now show canonical promotion previews for task/output drafts and let operators switch a persisted draft between a structured check and a review-only task-package scaffold
Persisted task/output drafts on cloud source pages now show what the current
promotion target will become in the canonical eval pipeline before anything
is saved.This makes the next step explicit: a draft can stay on the structured
deterministic path, or switch into a review-only task-package scaffold with the
expected environment and verifier placeholders called out up front.
CloudDashboard
Cloud source pages can now refine persisted task/output drafts through a review-only authoring agent, with an explicit fallback when Workers AI is unavailable
Persisted task/output drafts on cloud source pages can now run through a
bounded authoring-agent refinement step before the operator saves anything
canonical.This keeps the same draft/session contract, surfaces whether the refinement
came from a Think-backed path or a deterministic fallback, and lets the page
stay review-first even when Workers AI is unavailable. The dashboard exposes
that action through the new public refine route instead of a page-local-only
mutation path.
CloudDashboard
Cloud source pages now show review-only guidance for persisted task/output drafts and let operators apply that scaffold back into the draft before saving it
Persisted task/output drafts on cloud source pages now show API-derived
review-only guidance from matching suggestions, recent run pressure, and
saved eval-suite overlap.This means operators can see what failure a draft protects against, what
verifier shape is likely to fit best, whether they should extend an
existing suite, and apply the suggested scaffold back into the deterministic
draft before they save anything canonical.
CloudDashboard
Cloud task/output drafts now keep seed evidence, promotion target, and expected-outcome scaffolding, so deterministic eval authoring survives refreshes with real provenance instead of thin form state
Cloud source pages now persist richer deterministic draft metadata for eval
authoring, including the originating evidence, the intended promotion target,
and a first expected-outcome scaffold.This means review-only task/output drafts are no longer just local editor
fields. Operators can refresh and resume the draft while still seeing what
seeded it and what kind of deeper check it is trying to become.
CloudDashboard
Cloud source pages now keep one resumable task/output draft per source, so operators can refresh and continue deterministic eval authoring without rebuilding the draft from scratch
Cloud improve now stores a review-only task/output draft per source in the
runtime layer and exposes it back through source detail.This means a deterministic draft started from trigger evidence is no longer
purely local page state: operators can refresh, resume the draft, or clear it
without overwriting the saved trigger-confidence suite.
CloudDashboard
Improve-run detail now keeps the active timeline step and proposal links in sync after completion, so operators can see live progress and open the winning proposal without a manual refresh
Improve-run detail now seeds the current phase into the timeline immediately,
keeps the active step visibly live, and briefly re-checks proposal links
after a successful run until the winning proposal is available.This closes two gaps on the same surface: active runs now look active where the
operator is already reading, and completed runs no longer require a manual
refresh just to open the winning proposal.
CloudDashboard
Improve-run pages now keep refreshing proposal links briefly after a run completes, so new proposal links appear without a manual reload
Cloud improve-run pages now re-sync proposal links after terminal updates and
keep polling briefly when a winning candidate exists but the proposal link has
not settled yet.This closes the gap where a run could finish successfully and create a proposal
seconds later, while the improve page still looked like proposal creation had
been skipped until a manual refresh.
CloudDashboard
Cloud source pages now show live source-coordination state, including the active reserved run and queued rerun/cancel intent
Cloud source detail now carries a compact coordinator read model from the
runtime, and the source page surfaces that state directly.This makes it visible when a source already has an active reserved run, whether
cancel was requested, and whether a rerun is queued, without inferring it only
from raw run rows.
CloudDashboard
Cloud source pages can now promote trigger evidence into a detached task/output draft, so operators can scaffold deterministic checks without overwriting the saved trigger suite
Pending eval suggestions and recent run-pressure cards now offer a
Draft task check action that opens the eval editor in deterministic mode with a new,
unsaved task/output draft scaffold.This keeps the saved trigger-confidence suite intact while giving operators a
fast path to start deeper task/output coverage from real evidence.CloudDashboard
Improve-run timelines now keep the active phase on a colored dot, so the phase progression stays visible while a run is live
The live step on the cloud improve-run timeline now stays on the same
phase-colored dot as the rest of the run history instead of switching to a
generic spinner.This keeps the setup, evaluation, drafting, and finalization phases visually
distinct even while the run is still active.
CloudDashboard
Cloud skill pages now separate trigger-confidence evals from task/output checks in the onboarding and eval-editor language
Cloud skill onboarding and eval-editor copy now makes the intended eval
progression explicit: start with trigger confidence, then add task/output
checks once the skill is activating in the right places.This matches the hosted eval contract more closely and makes it clearer that
discoverability/routing checks and deeper task correctness checks answer
different questions.
CloudDashboard
Cloud source pages now expose a public rerun-analysis action, so operators can reprocess the current snapshot without creating a new upload or GitHub sync
Cloud source pages now include a
Rerun analysis action that triggers the
existing snapshot analysis pipeline for the current snapshot through a public
dashboard route.This makes it possible to refresh validation, lint, capability, and structural
reports after cloud-side analysis logic changes, without creating a new upload
or GitHub sync just to force a re-run.CloudDashboard
Improve-run pages now preserve run scope on per-candidate proposal links, so review navigation stays inside the run-specific proposal queue
Candidate-level
View proposal links on improve-run pages now keep the
current ?run=... context instead of dropping back to the global proposal
queue.This keeps the operator inside the run-specific review flow whether they open
the winning proposal from the summary header or from the candidate list.CloudDashboard
Proposal cards now render separate skill and review links instead of nesting anchors, fixing a hydration bug on the dashboard proposals page
Cloud proposal cards now show an explicit
Review proposal link inside the
card instead of wrapping the whole card in a detail link.This removes invalid nested anchor markup when a proposal card also links to its
skill page, which fixes the hydration error on the dashboard proposals route.CloudDashboard
Cloud source pages now read structural provenance directly from the proposal queue payload, removing an extra proposal-detail fetch from the latest-run review path
The run-scoped proposal queue now carries candidate provenance for cloud
improve proposals, so cloud source pages can explain the newest structural
proposal directly from the queue payload.This removes the extra proposal-detail request the source page used to make just
to show the latest proposal’s structural origin.
CloudDashboard
Cloud skill source pages now show which structural recommendation produced the newest linked proposal, so authors can connect snapshot analysis to the actual reviewable candidate
The structure-candidates panel on each cloud skill source page now surfaces
the newest linked proposal’s exact structural recommendation and deterministic
strategy when that proposal came from a structure run.This closes the gap between “top recommendations on the snapshot” and “the
candidate you can review right now,” so authors can see why the latest proposal
exists before opening the proposal detail page.
CloudDashboard
Cloud proposal queues and detail routes now have explicit review-flow coverage, reducing the risk of silent regressions in run-scoped proposal review
Added route-level coverage for the run-scoped proposal queue and proposal
detail lookup used by cloud-improve review flows.This hardens the path from improve-run pages into proposal review by checking
both the queue filter (
cloud_run_id) and proposal-detail validation behavior.CloudDashboard
Cloud improve and skills surfaces now use softer, consistent status treatments instead of mixing heavy cyan badges, raw enum labels, and duplicate readiness states
Cloud skill cards, improve-run pages, and setup checklists now use a
standardized status system with softer status chips, text-only status
labels where appropriate, and clearer warning/error semantics.This removes raw labels like
cloud_ready, collapses redundant Ready
states on the skills library cards, and brings advanced run settings into a
drawer with more readable model selection options.CloudDashboard
Proposal detail now normalizes structural provenance consistently after refresh and keeps run-scoped back navigation on the same helper path as the proposal queue
Proposal detail now uses shared helpers to normalize structural provenance and
run-scoped back navigation, so refreshed review pages keep the same
candidate-rationale and queue-return path instead of relying on ad hoc inline
logic.This keeps the proposal review surface more predictable as cloud-improve
candidates add richer provenance metadata and more scoped review flows.
CloudDashboard
Proposal detail now shows which structural recommendation produced a cloud structure candidate, so operators can see why a package rewrite was drafted before applying it
Cloud proposal detail now surfaces the structural recommendation and
deterministic strategy that produced a structure candidate, such as
extract_references or harden_script_ergonomics.That makes package-backed proposal review more defensible: operators can see
which structural analysis signal led to the candidate before deciding whether
to apply it to a draft or GitHub PR.CloudDashboard
Improve-run detail now links straight into the run-scoped proposal queue when a candidate frontier produced more than one reviewable proposal
Improve-run detail pages now preserve run context when you open the winning
proposal, and they surface a direct
Review run proposals action whenever a
hosted run produced more than one reviewable proposal.That keeps the run frontier, proposal queue, and individual proposal review
pages tied together, so multi-candidate review flows no longer require manual
navigation between /improve and /proposals.CloudDashboard
Run-scoped proposal review now preserves that scope when you open proposal detail, so it is easier to move through a single cloud-improve queue without losing context
When you open proposal detail from a run-scoped proposals list, the detail
page now keeps that run context and offers a
Back to run proposals path
instead of dropping you back into the global proposal backlog.This keeps multi-candidate cloud review flows tighter: operators can move
between the run-specific proposal queue and individual proposal detail pages
without reapplying filters or losing their place.CloudDashboard
Proposal review can now be scoped to a single improve run, so multi-candidate cloud runs no longer dump you back into the full proposal backlog
The proposals page now accepts a run-scoped view for cloud-improve runs,
letting operators review only the proposals created by a single hosted run
instead of filtering mentally through the full backlog.Cloud skill source pages now use that view when the latest run produced
multiple proposals, so the
Structure candidates panel can link directly
into proposal review even when there is more than one candidate to inspect.CloudDashboard
Cloud skill source pages now explain when structure proposals are viable and link directly into proposal review when the latest run produced one
Cloud skill source pages now synthesize the typed structural-analysis report
into a dedicated
Structure candidates panel instead of leaving operators to
infer package-shape readiness from raw technical details.The panel now shows whether the current snapshot is ready for structure
proposals, still review-first because of execution limits, or simply not a
structural change candidate right now. When the latest run already produced a
reviewable proposal, the page links straight into the proposal review flow;
otherwise it routes back to the latest run frontier.CloudDashboard
Cloud proposal review now renders the real candidate diff for package-structure changes and respects GitHub PR apply targets
Cloud proposal detail now renders the unified diff from the winning candidate
package instead of only showing placeholder
see candidate archive text.
Structure-focused proposals can be reviewed as real package changes, and the
page links directly to the candidate archive when you want the full package.Proposal apply now also respects the run’s configured apply target. If a
cloud-improve run was set to write back through GitHub, the proposal page now
offers a GitHub PR action instead of always defaulting to draft promotion.CloudDashboard
Cloud proposal detail now keeps the originating run and apply history visible after refresh
Cloud proposal detail now links back to the originating improve run so you
can move from a pending proposal into the full candidate frontier and
evaluation evidence without hunting for the run separately.The page also now shows recent apply attempts, including whether the attempt
targeted a draft or GitHub PR, when it ran, and the PR URL or error message
when one is available. That keeps proposal review useful even after the page
is refreshed or revisited later.
CloudDashboard
The proposals index now recognizes archive-backed cloud-improve proposals instead of showing them as fake single-field diffs
The proposals index now correctly tags cloud-improve proposals created by the
hosted runner, even when
proposed_by is the actual runner identity instead
of the older cloud_improve string.Archive-backed cloud proposals also no longer render (see candidate archive) as if it were a field-level diff. The card now tells you it is a
package-backed change and shows whether a linked candidate package and source
run are available before you open the full review page.CloudDashboard
Cloud skill source pages now surface structural analysis, script strategy, and frontmatter execution guidance instead of hiding them in raw validation blobs
Cloud source detail now renders the typed structural-analysis summary for the
current snapshot, including
SKILL.md line and token budgets, inferred
script strategy, compatibility notes, allowed-tools, and the execution flags
that explain whether cloud writeback is viable or still review-first.Validation checklists and technical details also now surface structural
recommendations as first-class findings instead of showing 0 issues for
typed reports that were previously stored outside the generic lint array
shape.CloudDashboard
Ended post-apply observation windows are now scored against baseline telemetry instead of staying pending forever
Applied cloud-improve proposals now compare a pre-apply telemetry window
against the finished post-apply observation window and classify the result as
helped, inconclusive, or regressed.Proposal detail now shows the before/after live-signal breakdown for eval
volume, pass rate, missed triggers, false negatives, and false positives, and
source-level Eval suite trust now folds recent observed regressions and
helps into the trust summary so coverage isn’t the only signal. Ops can
batch-score ended windows with
bun run score:cloud-improve-observations.CloudDashboard
Applied cloud-improve proposals now enter an explicit observation window instead of looking fully done the moment apply succeeds
Applied cloud-improve proposals now show a post-apply observation state on
the proposal detail page. Instead of treating
applied as the end of the
story, the page now distinguishes proposals that are still gathering live
signal from ones that will eventually be evaluated against observed outcomes.This is the first thin slice of the post-apply observation loop from the
cloud-improve quality hardening plan. It does not score outcomes yet, but it
does create a durable observation record the moment a draft promotion or
GitHub apply succeeds.CloudDashboard
Cloud skill pages now show eval-suite trust signals, and the cloud-improve runner has a first judge-calibration benchmark command
Cloud skill source pages now include an
Eval suite trust panel that shows
whether the saved suite still covers recent linked telemetry and the newest
hosted improve-run pressure. Instead of treating every eval win as equally
trustworthy, the page now marks suites as Fresh, Watch, Stale, or
No signal based on how much recent evidence is actually covered by saved
trigger-query cases.The cloud-improve runner also now ships with a first reusable
judge-calibration command, so the llm_judge trigger evaluator can be
checked against a labeled benchmark fixture instead of remaining an
unmeasured instrument.CloudDashboard
Cloud skill setup now auto-fixes common Agent Skills spec issues before the first snapshot is analyzed
Upload-backed and GitHub-backed cloud skill setup now canonicalize common
package issues before the first snapshot is analyzed. The setup flow rewrites
lowercase
skill.md to SKILL.md, rebuilds missing or invalid frontmatter,
normalizes the skill name to Agent Skills format, synthesizes a required
description when it is missing, and moves unsupported top-level frontmatter
fields into metadata so the package starts from a spec-compliant baseline.The setup response now also reports which fixes were applied, and the cloud
library success message calls that out before the first quick eval suite and
hosted improve run are prepared.CloudDashboard
Cloud eval suites now learn from telemetry and hosted runs, write accepted suggestions directly into saved suites, and surface eval pressure from the latest run
Cloud skill pages now surface suggested trigger-query cases from the linked
observed skill whenever recent real usage exposes misses or false positives
that are not already covered by the saved suite.Hosted improve runs also now persist query-level eval evidence, so recent run
failures and regressions can feed the same review queue when the active suite
no longer covers those cases.These suggestions stay review-first: you can append them into the draft suite
from the cloud page, inspect them in the editor, and then decide whether to
save them before the next improve run. Accepted and dismissed suggestions are
now also persisted, so the same pending cases do not keep resurfacing after
you review them. The cloud source page now also keeps a reviewed history with
restore and re-accept actions, so dismissed cases can be reopened and
accepted cases can be added back into the draft without losing their
provenance. Accepted suggestions can now also be written directly into the
selected saved suite with their telemetry/run provenance preserved, instead
of stopping at the draft-only state. The cloud source page now highlights the
latest run’s eval pressure directly, and the improve run detail page surfaces
the failed/regressed queries from that run instead of forcing operators to
dig through raw artifacts to find them. Source detail reads also now degrade
safely if the new eval-suggestion review tables are one migration behind,
while review actions return a clear migration-needed error instead of a raw
database exception.
CloudDashboard
Cloud improve now generates true bounded surface candidates and run pages render actual reviewed diffs
Hosted improve generation now treats
description, routing, and body
as real bounded mutation surfaces instead of prompt-only hints. A routing
candidate rewrites only routing, a description candidate rewrites only the
description, and body candidates preserve routing while updating the
non-routing sections they are allowed to touch.Improve run detail pages now also normalize old prose-only diff summaries
back into real unified diffs when the source and candidate archives exist, so
review pages show the actual changed lines instead of a rationale paragraph.CloudDashboard
Improve run detail pages now show the skill context, run outcome, and readable evidence instead of raw storage URLs
Hosted improve run pages now load the source skill and eval-suite context,
summarize the winning result or failure in plain language, and show the best
candidate’s score movement and diff preview directly on the page.Evidence is still available for download, but artifact links are now grouped
and labeled by what they represent instead of exposing a wall of raw R2
URLs.
CloudDashboard
Fresh cloud skills now auto-start the first hosted improve run, and the legacy in-process dispatcher is rollback-only
Fresh cloud skill sources now move directly from upload or sync into the
first hosted improve run once the quick eval suite is generated, instead of
stopping on the setup page and requiring a separate manual queue action.The API startup path also now treats the old in-process improve dispatcher as
an explicit legacy rollback path rather than part of the normal runtime.
Cloudflare remains the default hosted execution plane whenever the runtime
URL is configured.The API-key cloud-source surface also now matches the dashboard route for
creating GitHub-backed sources, which makes the same hosted improve flow
scriptable for smoke runs and automation.
New
@selftune/email package with 9 branded React Email templates: welcome,
alert notification, evolution proposal, weekly digest, team invitation, plan
upgrade, usage limit warning, getting started, and first insight. Alert emails
now use HTML templates instead of plain text. Team invitations and billing
checkout flows send branded emails automatically.Cloud
Cloud improve now auto-links uploaded and GitHub-backed sources to canonical skills, and imported task-package suites are first-class
Upload-backed and GitHub-backed cloud sources now automatically create or
reuse the canonical
skills row they belong to. That closes the proposal
gap where a winning improve run could persist artifacts but skip
proposal_created because the source had no linked skill_id.The eval-suite control plane also now accepts source_kind = imported for
deterministic task_package suites, which is the first explicit hosted lane
for benchmark-style imports instead of treating every imported suite as a
manual one. The docs now also include a first-class imported benchmark page
and script path for turning package manifests into live cloud eval suites.Cloud
Cloud improve now supports deterministic task-package eval suites and benchmark-style runtime docs
Hosted eval suites now accept deterministic
task_package cases, which lets
you point an improve run at a benchmark-style environment archive and verifier
script instead of relying only on trigger-query or exact-match checks.The Cloudflare runtime executes these task packages inside Sandboxes so the
verifier has a real filesystem and process boundary, and the public docs now
cover improve run events, statuses, and eval-suite API usage in the same
terminology the product uses.CloudDashboard
Improve run pages now show customer-facing live progress, delay states, and clearer timeline copy
Hosted improve pages now translate runtime activity into customer-facing
progress language instead of exposing queue, worker, or transport details.
Active runs surface clearer status cards, a friendlier timeline, and “taking
longer than expected” messaging when a run stalls.The improve overview also better distinguishes active versus completed work
without making the page feel like an internal operations console.
CloudDashboard
Improve run pages now refresh live while queued and running, with terminal refetch on completion
Hosted improve run detail pages now subscribe to the run event stream while a
run is
queued or running, updating phase and status live instead of
waiting for a manual refresh. When a terminal event arrives, the page
re-fetches full run detail so candidates, artifacts, and proposal state stay
in sync.The improve run list also now polls only while active runs are visible, which
keeps the overview current without constantly refetching completed history.Cloud
Cloud source uploads and GitHub sync now accept lowercase skill.md packages and preserve folder paths on the API-key surface
Hosted cloud-source ingest now accepts both
SKILL.md and lowercase
skill.md when validating uploaded packages and GitHub-backed skill repos.
That keeps upload and sync behavior aligned with the rest of the hosted
analysis pipeline, which already supported both casings.The API-key Hono upload route also now preserves multipart field keys as
relative paths instead of flattening uploaded files to their basenames, so
folder uploads keep nested references/ and other package structure intact.Added foundation for Cloudflare-backed improve run execution using Queues,
Workflows, and Sandboxes. A new
GET /api/v1/improve-runs/:id/events endpoint
streams run lifecycle events via SSE, enabling live updates on the run detail
page without manual refresh. The runtime mode is controlled by
CLOUD_IMPROVE_RUNTIME_MODE and defaults to legacy with no behavior change
until explicitly switched.CloudDashboard
Cloud skill validation now uses native spec checks and clearer report detail during setup
Cloud skill setup now persists one validation report per snapshot and shows
those results inline in the guided setup hero, so structural validation,
best-practice lint, and capability classification are easier to inspect
without dropping into raw logs.The hosted validation step also now runs on a native TypeScript
implementation of the Agent Skills frontmatter rules instead of shelling out
to the demonstration
skills-ref toolchain. That keeps cloud validation
deterministic in production while tightening allowed-tools parsing and
preserving clearer per-rule issue messages.Apply flows also now version and re-upload promoted skill archives more
safely. Draft apply and GitHub PR apply both keep the cloud source pointed at
the newly promoted snapshot, preserve archive manifests across lowercase
skill.md packages, and avoid corrupting frontmatter when YAML values
contain ---.CloudDashboard
Cloud source APIs now honor source-type and skill filters consistently across dashboard and API-key surfaces
The hosted cloud-source list API now applies the same
type and skill_id
filters on the API-key Hono surface that the dashboard session route already
supported. That keeps browser, CLI, and smoke-test callers on one normalized
contract when listing cloud skills.This batch also tightens the hosted improve apply/runtime path so root-level
GitHub applies do not infer repo-wide deletions, runner dependencies resolve
snapshots through the correct org-scoped database client, and local Neon CLI
binding metadata is no longer tracked in git.CloudDashboard
Cloud improve model selectors now load the live OpenRouter catalog and use a teacher-student default spread
The per-run model selectors on cloud and observed skill detail pages no
longer use a hardcoded GPT-only shortlist. They now load the current
OpenRouter model catalog from the server and expose the broader set of
text-capable models available through the hosted cloud runtime.Each selector also now carries an explicit recommended default for
generate, judge, and summarize. Leaving a selector empty keeps the
server-side default for that role, and the UI now spells out those defaults
directly so you can test alternatives without losing track of the intended
baseline. The selectors are now searchable comboboxes as well, so longer
model lists stay usable without scrolling through a giant dropdown.The recommended defaults now follow a clearer teacher-student split
instead of a flat GPT-only stack:
google/gemini-2.5-pro for proposal
generation, google/gemini-2.5-flash for judging, and
google/gemini-2.5-flash-lite for summarization. That keeps the strongest
model on the expensive generative step while moving validation and helper
work onto cheaper OpenRouter models.CloudDashboard
Cloud skill onboarding now auto-creates a 50-case eval suite and funnels into one happy path
When you create or sync a cloud skill, the detail page now automatically
drafts and saves a 50-case quick eval suite from the current snapshot
instead of making you build one manually first. The skill page now leads with
a single guided decision: edit the generated eval suite or start the
hosted improve run.The cloud skill detail UI has also been simplified around that progression.
The primary suite is now treated as one editable artifact with save support,
advanced run controls stay collapsed by default, and metadata/report panels
are moved behind a technical-details disclosure so the page feels less like a
control plane and more like a clear product flow.The Eval summary card now surfaces the per-case origin (auto-generated
versus hand-curated versus a mixed breakdown) so you can tell at a glance
whether the suite is still the synthetic draft or has been edited. Clicking
Edit eval suite also now smooth-scrolls the editor into view, and the
hero action row has been reordered so advanced run options live next to the
edit affordance rather than after Start improve run.
CloudDashboard
Overview now presents the hosted cloud loop instead of legacy first-run telemetry onboarding
The cloud dashboard overview now introduces SelfTune as a hosted review
loop instead of the older “run selftune and wait for skills to appear”
onboarding. The empty-state banner now points people toward the real cloud
path: create or import a cloud skill, shape a reviewable eval suite, run the
hosted comparison, and review the resulting proposal before draft apply.The overview also now keys that banner off cloud authoring state rather
than observed telemetry alone, so the first-run guidance stays visible until
you actually have cloud sources in the hosted product.
The cloud skill detail page now includes a real Quick Eval Suite editor
instead of only raw textareas and JSON authoring. Trigger-query cases are now
editable as table rows with expectation, invocation type, provenance, and
row-level remove actions.For
llm_judge suites, the page can also draft a synthetic seed directly
from the current snapshot’s SKILL.md. Those seeded cases are marked as
synthetic so you can review, revise, or delete them before creating the
hosted eval suite.CloudDashboard
Cloud folder uploads now preserve a real skill directory and show clearer selection state
Cloud skill uploads now package the selected folder as a real skill directory
instead of flattening only its file contents into the snapshot archive. That
means downstream validation sees the skill in a proper directory layout,
which fixes false failures caused by generic wrapper names during
skills-ref
analysis.The dashboard upload flow is also clearer: the picker now behaves like a
folder intake card, shows the detected folder name and file count, confirms
whether a root SKILL.md was found, and disables upload until the selection
is actually valid.CloudDashboard
Cloud improve runs now support separate generate, judge, and summarize model overrides
Cloud improve run setup now exposes three separate model selectors instead of
one shared override. You can independently pick the model used for
candidate generation, LLM judging, and summarization from the
skill detail page before queueing a run.These overrides are stored with the run itself and passed through the hosted
runner, so they no longer collapse into a single model choice. This makes it
practical to test cheaper summarize settings while keeping a stronger judge,
or to isolate generation changes without touching the server defaults.
CloudDashboard
Cloud dashboard now uses skill-first wording instead of exposing raw source terminology
The cloud dashboard now uses skill-first wording across the library,
detail pages, and observed-to-cloud bridge instead of exposing the backend
source model directly in the UI. Cloud library cards, blocked-state
messages, quick eval setup, and linked-skill surfaces now read as normal
product concepts: Cloud skills, Open Skill, and Linked cloud
skills.This is a terminology cleanup only. The backend cloud_skill_sources model
and API routes are unchanged, but the visible dashboard flow is less
confusing because it no longer asks users to think in storage-layer terms.Observed skill cards now include an Import to Cloud button that navigates
directly to the cloud library with the skill pre-filled for import. The cloud
source detail page also gains an Eval Suites section where you can view
existing suites scoped to a source and create new ones inline with a name,
verifier kind, and JSON test cases — no CLI or API calls required.Proposal detail pages now show full eval comparison data, artifact kinds,
confidence levels, and an Apply to Draft button that closes the loop
from review to draft apply. The jobs page shows Cloud Improve Runs
alongside pipeline jobs with status badges and candidate counts.
CloudDashboard
Cloud Library now shows cloud authoring records separately from observed telemetry skills
The cloud dashboard’s main
Skills surface now reflects the real cloud
authoring model instead of the telemetry-backed skill table. It lists
GitHub-backed sources, imported uploads, and cloud-managed records with their
current snapshot and capability state.The old telemetry-backed skills library is still available, but it now lives
under Observed so local and cloud concepts do not get mixed together. Cloud
sources without a linked telemetry skill now have their own detail page for
snapshot metadata, validation reports, and hosted improve controls.The cloud library also now includes first-run onboarding directly in the UI:
you can upload a skill folder into a new cloud source, create and sync a
GitHub-backed source from a bound installation, jump into cloud import from
observed skills, and create a lightweight eval suite from the cloud detail
page before queuing a hosted improve run.Observed skill detail pages now also expose that bridge directly: if a skill
already has linked cloud sources you can jump straight into them, and if it
does not you can import it into cloud from the report itself instead of
backing out to the library first.CloudDashboard
Cloud improve runs can now override the model per run, with cheaper default summarize policy
Cloud improve runs can now override the hosted model policy directly from the
skill page before queueing a run. The selected override applies to the full
run, not just candidate generation, so you can force a cheaper test model or
a stronger one without changing server env defaults.Hosted defaults are also now more cost-aware for testing: generation and
judging stay on
openai/gpt-4.1-mini, while summarize and low-risk helper
work default to openai/gpt-4.1-nano.Cloud
Cloud skill improvement integration: runner package, eval backends, control-plane wiring, and hosted eval-suite parity
The cloud skill improvement pipeline is now fully wired end-to-end. The
isolated runner package (
@selftune/cloud-improve-runner) connects to the
control-plane orchestrator via concrete dependency adapters. Eval backends
(trigger-query LLM judge + deterministic) dispatch through a registry. Both
draft and GitHub apply paths consume the same candidate archive contract.Hosted eval-suite creation now validates runnable manual suites for both
llm_judge and deterministic verifiers through the same control-plane
contract the runner consumes, which keeps the dashboard and API-key paths on
one canonical suite definition.selftune improveandselftune evolvenow fall back to plain stderr progress lines when the terminal is not a TTY, instead of going completely silent while long proposal or validation steps are still running. - Interactive terminals keep the spinner/TUI behavior, while test runs remain quiet by default.
- Local dashboard action toasts now include a
Live runaction that opens the exact/live-runentry for the streaming creator-loop event, including the event id, skill, and action selection state. - The floatingLive lifecycle actionsfeed now uses the same deep link, so clicking a running or finished lifecycle card jumps straight into the matchingLive Runentry instead of leaving you to find it manually.
selftune eval generatenow accepts--agentfor--synthetic,--auto-synthetic, and--blend, so you can forceopencode,codex, orpiinstead of relying on auto-detection order. - Cold-start synthetic eval generation now reuses the same cleaned query filtering as log-derived evals and summarizes oversizedSKILL.mdcontent before sending it to the runtime, which reduces prompt bloat for large skills likeSelfTuneBlog.
- Bounded package search now writes merged routing/body candidates into a new
temp package snapshot instead of overwriting the already-evaluated body
variant on disk, so candidate artifacts remain consistent for later winner
application and review. -
selftune create publish --watch --ignore-watch-alertsnow also bypasses the watch gate when the watch subprocess crashes or fails to emit structured JSON, while still surfacing the warning and remediation command.
- The OSS local dashboard
LiveRuntest fixture now uses the realDashboardActionResultSummaryshape for bounded package-search summaries, so export verification no longer fails whensearch_runis present on deploy candidate entries.
- The local dashboard now normalizes
selftune create replay,selftune create baseline,selftune evolve,selftune evolve body, andselftune search-runinto the same lifecycle-facing commands the CLI already shows, so Overview, Skill Report, and Live Run no longer leak stage-level command names for draft-package flows.
selftune create baseline --mode packagenow reuses the last fresh with-skill replay from the canonical package-evaluation artifact when the draft fingerprint still matches, so measuring baseline no longer pays for two full replay passes after an unchangedverify,report, orsearch-run. - Package baseline now emits explicitwith_skill_replayandwithout_skill_replaystep progress so the local dashboard live-run surface shows immediate movement instead of looking stuck while the underlying replay work is still running.
selftune search-runnow prefers reflective routing/body proposals from measured runtime failures before targeted or deterministic fallback. - When routing and body both produce accepted improvements, package search now evaluates a merged candidate before final winner selection instead of forcing the frontier to choose between complementary single-surface edits. - Plainselftune improvenow auto-selects bounded package search for skills that already have package evidence or a draft package manifest, so agents do not need to force--scope packagefor the main package-shaped lifecycle. - Added an end-to-end package lifecycle test coveringverifyauto-fix, bounded package search, winner promotion, andpublish --watch.
selftune search-runnow uses the same measured targeted-routing/body mutation path as orchestrate package search, falling back to deterministic variants only when targeted variants do not fill the requested minibatch. - The public CLI docs and workflow docs now describesearch-runas bounded local package search over draft variants instead of registry lookup, and theEvolveworkflow now points package-scope users at the measured targeted search path instead of only the older deterministic description. - Publish/package-search lifecycle docs now describe the real blocking publish-time watch gate instead of the older advisory wording.
selftune verifynow auto-runs the real missing-evidence commands with the required flags and skill context, including--auto-syntheticeval generation and generated unit tests. -selftune create publish --watchnow blocks publish if the watch subprocess fails or returns malformed output instead of treating missing watch JSON as a passing gate. - Eval-informed targeted mutations now readgrading_results.pass_rate,expectations_json, andfailure_feedback_jsonfrom the real SQLite schema instead of a test-onlysummary_jsonshape. - The shipped lifecycle docs now describe the actual concrete readiness states and the correct--ignore-watch-alertsflag.
normalizeLifecycleCommandnow mapscreate replay,create baseline,evolve,evolve-body, andsearch-runto their lifecycle equivalents. -selftune --helpnow shows Primary Lifecycle commands first, with Advanced / Stage Commands below.
selftune verifyauto-runs missing evidence steps (up to 4 iterations) when readiness checks fail. Use--no-auto-fixto skip.
collectPackageSearchEligibleSkillsnow includes a second eligibility tier: skills with aselftune.create.jsondraft package and at least 3 grading results in the DB are routed to package search during orchestrate. - The existing frontier/artifact fast path is unchanged; the new tier is additive and fail-open (skips silently if the grading table is missing).
SearchRun.mdno longer claims orchestrate cannot auto-select package search — it documents the eligibility criteria and plan-phase routing. -Watch.mdadds a “How Watch Evidence Feeds Back to the Frontier” section explaining watch rank levels, SQLite row updates, and dashboard visibility. -SKILL.mdSearchRun routing keywords now include “optimize package”, “improve routing and body together”, and “bounded evolution”.
create publish --watchnow blocks publishing when the watch gate detects active alerts (published: false,watch_gate_blocked: true), instead of unconditionally publishing. Use--ignore-watch-alertsto bypass. -extractMutationWeaknessesnow populatesgradingFailurePatternsfrom theexpectationsarray in grading summary JSON, enabling targeted body mutations to focus on specific failed expectations.
- Orchestrate now marks skills package-search-eligible from the real accepted frontier and canonical package-evaluation artifacts, so the new package-search branch is reachable in normal runs instead of existing only in isolated tests.
- The orchestrate package-search phase now uses the current mutation and
winner-application contracts, including targeted routing/body variants,
current candidate path fields, and the current
applySearchRunWinnerresponse shape. -create publish --watchnow surfaceswatch_gate_passed,watch_gate_warnings, andwatch_trust_scoredirectly in the publish payload, and--ignore-watch-alertsnow intentionally bypasses that advisory gate when needed. - Skill reports now populatewatch_trust_scorefrom the latest stored package-evaluation watch summary, so the dashboard watch trust indicator renders from real watch evidence instead of staying empty. - Fixed theselftune orchestrateCLI docs page so Mintlify renders it as a normal document instead of a raw fenced code block. - Dashboard skill report and live run now display routing and body weakness percentages from surface plan data, with a visual bar highlighting the weaker surface. The frontier panel also shows a parent-vs-winner comparison when both members are available.
- Added evidence-driven scope selection to orchestrate so it automatically chooses between description-level evolve and package-level bounded search based on accepted frontier state and canonical package evaluation evidence. - Added watch trust scoring feedback so post-deploy regressions can demote accepted frontier candidates and influence future scope selection. - Updated workflow and skill documentation to reflect the new package-search-in-orchestrate truth.
- Added deterministic routing mutations (synonym expansion, granularity split, coverage broadening) and body mutations (instruction emphasis, example enrichment, description expansion) for bounded package evolution. - Added eval-informed targeted mutations that consume measured weaknesses from replay failures and grading results to focus routing and body changes on specific failure patterns. - Added weakness extraction from the local SQLite database to surface replay failure samples, routing misses, body quality scores, and grading pass rate deltas for mutation targeting.
- Added
computeWatchTrustScoreto the watch module, producing a 0-1 trust score from trigger regression, grade regression, and rollback signals. - Added an advisory publish watch gate that warns when active alerts or low trust scores are detected, with--ignore-watch-alertsbypass for experts. - Extended the dashboard contract withwatch_trust_scoreon skill reports andwatch_gate_passedon action result summaries. - Updated the live run screen to display watch gate pass/alert badges when watch or deploy actions complete. - Added a watch trust indicator to the skill report creator loop section.
- Added a package-search candidate action to the orchestrate loop so skills with accepted package frontier candidates are routed through bounded package search instead of standard evolution. - The new phase generates bounded mutations, fingerprints variants, runs package search evaluation, and applies winning candidates automatically. - Package search modules are lazy-loaded and gracefully degrade when unavailable, so the existing orchestrate flow is unaffected until the full package search stack is present.
search-runnow treatsbody.quality_score: nullas a neutral weakness signal when the body already passed validation, instead of coercing it to maximum weakness. - This prevents--surface bothfrom over-allocating routing/body search budget toward body mutations when quality assessment was unavailable but the current body was still valid.
search-run --surface bothnow reads the accepted frontier first and falls back to the canonical package evaluation when needed, using that measured package state to bias routing/body candidate counts. - This replaces the old fixed half-routing half-body split with a weakness planner that sends more of the minibatch budget toward the weaker measured surface while still keeping bounded deterministic search behavior. - The chosen surface budget is now persisted into search provenance and shown in the live-run and skill-report frontier surfaces, so reviewers can see why a run spent more budget on routing or body.
- Added
search-run --apply-winner, which copies the winning candidate back into the draft package and refreshes the canonical package-evaluation artifact from the accepted candidate cache instead of leaving search as read-only provenance. -selftune improve --scope packagenow adds winner promotion by default and keeps--dry-runas the review-only escape hatch. - Search-run dashboard summaries now carry the resulting next command and package-evaluation context when a winning candidate is applied, so live review stays grounded in measured package state instead of raw search provenance.
- Added
selftune improve --scope package, which routes the primary improvement alias intoselftune search-runinstead of keeping bounded package search behind an expert-only command. - Package scope now preserves--eval-set, strips redundant--dry-run, normalizes compatible replay validation flags, and maps--candidatesontosearch-run’s--max-candidatesknob. - Updated command help, workflow docs, SKILL routing guidance, and CLI docs so package search is taught as part of the main measured improvement loop.
- Added
selftune search-runas a real top-level CLI command that generates bounded routing/body package variants, evaluates them through the shared package evaluator, and persists the selected winner plus provenance. - Wiredsearch-runthrough dashboard actions, child-process event instrumentation, live-run summaries, and draft-package action buttons so bounded search is executable from the product surface instead of only existing as stored backend state. - The skill report backend now returns real package frontier state and the latest search-run provenance, so the frontier panel is driven by measured candidate history rather than a dormant response field. - Package search evaluations now normalize temp candidate variants back onto the canonical skill name, and winner selection now follows the accepted frontier over the full evaluator contract instead of replay-only gains. - Updated command help, workflow docs, SKILL routing, and the CLI quick reference so the new search surface is documented consistently.
- Updated
selftune statusoutput to label the readiness section “Package pipeline” instead of “Creator loop”. - Adapted package search runner to the mature evaluator API with frontier-based parent selection. - Normalized SKILL.md description and body to reference the package evaluation pipeline (replay, baseline, grading, body, unit tests, and post-deploy watch) as the primary improvement mechanism. - Updated Evolve, EvolveBody, Watch, and CreateTestDeploy workflow docs to use package evaluation pipeline terminology consistently. - Normalized Baseline, Evals, UnitTest, SignalsDashboard workflow docs and creator-playbook reference to use package evaluation pipeline terminology.
- Added package frontier panel to skill report showing accepted candidates ranked by measured evidence with watch-fed demotion indicators. - Added search run panel to live run screen showing selected parent, candidates evaluated, winner determination, and provenance detail. - Added search-run action result parsing to the dashboard action result contract so search runs surface structured summaries alongside existing replay dry-run results.
- Added
generateRoutingMutations()andgenerateBodyMutations()in the evolution pipeline to produce complete skill file variants that a package search runner can score. Three routing strategies (synonym expansion, granularity split, coverage broadening) and three body strategies (instruction emphasis, example enrichment, description expansion) create bounded variants written to temporary directories.
- Added bounded package search runner that evaluates candidate skill variants against the accepted frontier parent with measured delta acceptance. - Added package candidate state management with frontier reading, parent selection, and fingerprint-based deduplication. - Added package search provenance persistence tracking frontier size, parent selection method, candidate fingerprints, and evaluation summaries.
selftune watchnow reads the current package-evaluation artifact when one exists and computes an efficiency regression signal from observed post-deploy sessions, instead of only looking for trigger-pass-rate regressions and optional grade regressions. - Efficiency watch is grounded in measured package baselines already produced bycreate reportandcreate publish, so post-deploy monitoring now compares observed input tokens, output tokens, and assistant turns against the same package-evaluator contract used before publish. - Efficiency regressions now flow through the structured watch result and the nested package watch summary, so publish/watch consumers can surface the same measured signal without scraping alert text. - The local dashboard watch parser now preserves those efficiency-regression fields in the package watch summary, keeping the watch contract forward-ready for richer live-run presentation as more post-deploy package signals land.
- Durable draft package candidates now carry a measured acceptance decision in
local state, instead of only lineage metadata, so candidate history can
distinguish accepted improvements from measured regressions. - Acceptance is
computed from package-evaluator evidence rather than model confidence, with
explicit replay, routing, baseline-lift, body-quality, and unit-test deltas
plus a human-readable rationale attached to the candidate summary. -
Re-evaluating the same draft fingerprint preserves the original parent
relationship instead of inventing a new comparison target, so repeated review
runs update the candidate record without corrupting lineage. - Fresh
candidates now compare their measured acceptance against the latest accepted
frontier member instead of blindly inheriting the most recent rejected draft
as the comparison baseline, while still keeping chronological lineage in the
parent link. - When the current draft matches an already accepted frontier
member, package evaluation can now reuse that candidate-specific artifact by
fingerprint even if the canonical latest package report points at some other
draft, so re-checking an accepted draft no longer repays the full evaluator
cost. - Accepted-frontier selection is now ranked by measured package outcomes
instead of timestamp alone, so newer accepted drafts with weaker grading or
weaker observed health no longer automatically become the comparison parent
for the next candidate. -
create publish --watchnow writes structured watch results back into the matching package candidate artifact and registry row, so observed regressions can demote an accepted draft in later frontier selection without fabricating a brand-new evaluation event. - Cached package-evaluation reuse now also requires acceptance metadata in the stored artifact, so older lineage-only artifacts automatically refresh once before they can participate in candidate-aware reuse. - Benchmark reports,create publishsummaries, and the local dashboard live-run screen now surface the candidate acceptance decision and rationale, so measured accept/reject state is visible without opening archived JSON.
- Fresh draft package evaluations now register a durable package candidate per package fingerprint in local state, instead of only overwriting one latest package report per skill. - New candidate records carry parent linkage to the previously evaluated draft for the same skill plus a candidate-specific archived evaluation artifact, so later bounded package search can reuse lineage and evaluator evidence instead of rebuilding history from ad hoc files. - Cached package-evaluation reuse now requires candidate metadata in the saved artifact too, so older artifacts automatically force one fresh measured run before they can participate in candidate-aware reuse. - Benchmark reports, publish summaries, and the local dashboard live-run view now surface candidate ID, parent linkage, and generation directly, so candidate lineage is inspectable without opening archived JSON artifacts.
- Repositioned the shipped
selftuneskill around a smaller lifecycle:Create,Verify,Publish,Improve, andRun, instead of leading with the older stage-heavy creator loop. - Added new primary workflow docs forVerify,Publish,Improve, andRun, while keeping the existing lower-level eval, replay, baseline, watch, and body-evolution workflows available as advanced surfaces. - UpdatedSKILL.md, routing keywords, and lifecycle-state guidance so “can I trust this skill?”, “ship this skill”, and “run the loop” now map to intention-level workflows that still use today’s commands accurately under the hood. - ReframedCreateas draft authoring only, marked the olderCreateTestDeployworkflow as legacy compatibility guidance, and taughtOrchestrateas the underlying runtime behind the simplerRunconcept. - The local dashboard action stream and dashboard-triggered publish/evolve paths now recognize and use the newverify,publish,improve, andrunaliases where they preserve the same measured behavior, so the live-run UI stays aligned with the simplified lifecycle surface. - The local dashboard overview, skill report, live action feed, and CLI docs now teach draft-package work asverify,publish, and live monitoring first, while still exposing the lower-level eval, replay, baseline, and create-check commands when an agent needs to drive the advanced loop manually. -selftune status, dashboard recommended commands, live-run next-command cards, the shipped quick reference, README, and the main skill-authoring guides now normalize old surface aliases likecreate check,create publish, andorchestrateintoverify,publish, andrunwhen the underlying behavior is equivalent, so the product stops teaching mixed lifecycle vocabulary by default. - Scheduled automation surfaces now teachselftune runas the default autonomous loop entrypoint: cron job messages, generated schedule snippets, alpha-enrollment guidance, orchestration reports, and the related docs and skill workflows all userunfirst while keepingorchestrateas the underlying advanced runtime name where needed. - Fixed theselftune createCLI page after a broken MDX wrapper landed, and updated the main authoring, troubleshooting, sharing, trigger-testing, and creator-playbook docs so they teachverify/publishfirst while still documenting the lower-levelcreate replay/create baselinepackage steps when a draft needs explicit measured proof. - Normalized the secondary advanced workflow docs and README soeval,unit-test,baseline,evolve,evolve body, dashboard live-run, and legacy create-test-deploy guidance now distinguish draft-package lifecycle work from already-published skill iteration, instead of re-teaching the old creator-loop chain as the default. - Cleaned up the remaining lifecycle wording instatus,eval, andcreateCLI docs plus the shippedSKILL.mdreference table, so “creator loop” now mainly survives as a compatibility/search term instead of the default label for the product surface. - Corrected the package-search docs sosearch-runandimprove --scope packageare documented as explicit bounded-search surfaces, without claiming thatrun/orchestratealready auto-select package search before that automation is actually shipped.
- Added a canonical full-evaluation artifact beside the stored package
summary, so
create reportand publish-time package gates can reuse one measured replay/baseline/body-validation result instead of scraping or recomputing partial state. - Package-evaluation reuse is guarded by the bounded package fingerprint and request shape, so edited drafts or changed evaluation requests still trigger a fresh measured run instead of trusting stale evidence. - Cache hits only apply when the saved package artifact already includes the current routing/body validation dimensions, so older summaries automatically fall back to a fresh measured run instead of silently downgrading the review signal. - Benchmark reports and publish output now label whether the package evaluation was freshly measured or reused from a matching artifact cache, so creators can audit reuse instead of inferring it from timing or logs. - The local dashboard live-run summary now surfaces that same fresh-vs-cached evaluation source for package report/publish actions, so cache reuse stays visible in the main review UI too.
- Extended the shared draft package evaluator so
create reportandcreate publishnow attach current routing replay validation and current body validation alongside replay, baseline, grading, unit-test, and watch evidence. - Updated the benchmark-style package report format so routing replay and body validation show up in the same deterministic artifact as the rest of the measured package evidence. - Updated the active bounded package-evolution plan to reflect that body/routing validation is now part of the unified evaluator contract, moving the remaining gap toward candidate state, evaluator reuse, and measured search rather than missing evaluator dimensions.
- Added
selftune create report --skill-path <path>as a no-side-effect package-evaluation command that runs replay plus baseline and renders one benchmark-style report with failure analysis, measured lift, recommendation, and next-step guidance. - Added the same report shape as a reusable helper in the shared draft package evaluator so future dashboard and PR-summary surfaces can reuse one deterministic evidence format instead of inventing ad hoc summaries.
- Updated the selftune skill workflow docs, quick reference,
README, and CLI docs so package creators can explicitly request a measured
publish-readiness report before running
create publish.
- Updated
selftune create publishso draft-package publishing now re-runscreate replay --mode packageandcreate baseline --mode packageas the final measured gate before watch. - Removed the old direct handoff fromcreate publishinto description-onlyselftune evolve, keeping the creator loop grounded in package-level validation instead of a description mutation step. - Added a shared package-evaluation summary thatcreate publishcan return directly, so draft deploy/watch actions have one measured result shape instead of stitching together replay and baseline outcomes ad hoc. - Updated the local dashboard action parser so draft-package baseline and publish runs can surface replay mode, before/after pass rates, and lift on the live run screen. -selftune watchnow emits a machine-readablerecommended_command, andcreate publish --watchnow carries the nestedwatch_resultpayload through directly so draft publish/watch flows expose measured post-deploy pass rates, alerts, and rollback recommendations instead of only a coarse “watch started” status. - Updated creator-loop readiness andselftune statusguidance so draft packages now recommendcreate replay,create baseline, andcreate publishinstead of falling back to the olderevolve/gradecommands for those milestones. - Updated the overview, skill report, andselftune statuscreator-loop surfaces so draft packages stay blocked oncreate checkor package-resource fixes until those checks actually pass, instead of skipping ahead to replay or publish because later creator-loop artifacts already exist. - Added dashboard support forcreate checkas a runnable draft-package action, so the live-run screen and draft package panel can stream and summarize spec-validation checks instead of showing that step as copy-only guidance. - Added structured progress events forcreate check, so the live-run screen now shows draft-package load, Agent Skills validation, and selftune readiness computation as explicit steps instead of only the final JSON result. - Made the overview creator-loop priorities runnable from the dashboard for actionable steps, so top-level draft-package cards can launchcreate check, eval generation, replay, baseline, and publish flows without drilling into the per-skill report first. - Updated the CLI help, OSS workflow docs, and docs site reference so the publish contract matches the package-first creator loop. - The live-run summary tiles now relabel watch actions asBaseline,Observed,Delta, andSignal, so post-deploy watch evidence no longer appears under the older dry-runBefore/After/Validationvocabulary. - The shared package-evaluation payload now carries runtime efficiency and representative evidence, so package replay / baseline / publish flows can return measured duration and token aggregates together with replay-failure and baseline-win samples instead of only pass-rate summaries. - The live-run screen now surfaces those measured package-evaluation artifacts directly, including replay-failure samples, baseline-win/regression samples, with-skill versus without-skill efficiency totals, and recommended next commands when publish or watch actions expose them.
- Added
report-packageas a first-class dashboard action for draft skills, so the skill report and live-run feed can launchselftune create reportdirectly and label the resulting benchmark artifact separately from baseline, publish, and watch runs. create publish --watchnow attaches a structured watch summary to that same package-evaluation payload, and the live-run screen renders watch snapshot counts, invocation-type totals, rollback state, and grade-watch deltas from that shared measured contract.- Clarified the public CLI docs and shipped
Createworkflow so agents can rely on both the raw nestedwatch_resultpayload and the normalizedpackage_evaluation.watchblock when they parse publish-with-watch results. selftune evolveandselftune evolve bodyno longer reject proposals before measured validation solely because model-reported confidence is low;--confidencenow acts as a review threshold and adaptive-gate risk signal instead of a hard pre-validation stop.- The shared package-evaluation payload now also includes grading baseline
versus recent grading deltas when that data exists, so
create reportandcreate publish --jsoncan show observed execution-quality movement next to replay, baseline, and watch evidence. - The local dashboard now parses and renders that same
package_evaluation.gradingblock in live-run summaries, so draft package report and publish flows expose measured grading movement without requiring raw JSON inspection. - The latest package-evaluation summary is now stored canonically in SQLite
and mirrored to
~/.selftune/package-evaluations/<skill>.json, so draft report/publish/watch flows can reuse one measured artifact instead of treating package evaluation as stdout-only output. - Draft-package readiness and
create checknow honor the latest stored package-evaluation status, so a measuredreplay_failedorbaseline_failedresult keeps the skill blocked on the corresponding package gate instead of surfacing a falseready to publishstate just because the older replay or baseline artifacts exist. - The shared package-evaluation payload now also carries deterministic unit
test results and representative failing tests when that evidence exists, so
create report,create publish --json, and the live-run UI can review the latest measured test run alongside replay, baseline, grading, and watch evidence. - Draft-package readiness and
create checknow also honor the latest failed deterministic unit-test run when one exists, so stored test failures keep the draft blocked on rerunning unit tests instead of treating test-file presence alone as publish-ready proof. - Stored package-evaluation artifacts now include a bounded package fingerprint, and draft-package readiness only trusts those replay/baseline results when the fingerprint still matches the current package tree, so stale failed measurements stop blocking edited drafts just because they share the same skill name.
- Fixed dashboard child-process action context for
report-package, socreate reportandverifynow stream live progress and metrics events into the live-run screen instead of silently dropping them when the action context is read from environment variables.
- Added
selftune create initas the clean-slate authoring path for new skills. - Addedselftune create scaffold --from-workflow ...as the workflow-derived authoring path, and upgradedselftune workflows scaffoldto emit the same package shape for backward compatibility. - Package drafts now includeSKILL.md,workflows/default.md,references/overview.md, emptyscripts/andassets/directories, plus aselftune.create.jsonmanifest. - Added
selftune create checkto run Agent Skills spec validation first and then compute selftune-specific package readiness for evals, unit tests, replay, and baseline. - Addedselftune create replay,selftune create baseline,selftune create status, andselftune create publishso the draft-package path now reaches all the way through replay validation, lift measurement, and handoff into the existing evolve/watch surfaces. - Added package-mode replay staging so runtime replay can read workflow/reference files inside the staged skill package without treating them as unrelated paths. - The local dashboard now surfaces draft packages before they have live telemetry, shows package-local create readiness on the skill report, and routes dashboard replay/baseline/ publish actions through the draft-aware create commands automatically. -selftune create checknow recommendscreate replay,create baseline, andcreate publishfor draft-package next steps instead of the older generic evolve/grade commands, keeping package-tree staging consistent from CLI output through the dashboard. - Hardened the local dashboard draft-package views so the exported OSS app typechecks cleanly when create-readiness data is optional, preserving the draft-package panels in shipped builds. - Fixedselftune workflows scaffold --writeso fresh workflow-derived packages are written through the shared draft-package writer instead of pre-creating the directory and tripping the overwrite guard. - Draft-package dashboard actions now start eval generation with--auto-synthetic, so cold-start skills can bootstrap eval sets from the dashboard instead of attempting empty log-based generation. - Added agent workflow docs and public CLI docs so agents can route package authoring requests to the full command surface.
- Added GitHub App installation binding for cloud orgs so a team can associate
a GitHub installation with its registry workspace. - Added GitHub-backed
registry connection APIs for listing accessible repos, connecting a repo to a
registry entry, disconnecting it, and requesting manual sync. - Added
immediate manual sync publishing so a connected repo path is packaged from
GitHub, archived, and pushed into the registry as a GitHub-sourced version
without waiting on a background worker. - Added webhook-driven auto-publish
for default-branch pushes and matching Git tags so connected repos now flow
into the registry without manual sync. - Added a dashboard GitHub settings
flow with installation binding, repo discovery, monorepo path selection, and
connection management controls. - Added Tier A GitHub write-back with
org-level policy, per-connection opt-in, persisted publish attempts, and
optional commit status/check-run updates for successful, skipped, and failed
publishes. - Added direct
selftune registry install github:owner/repo[@ref][//path]support so skills can be installed straight from GitHub with monorepo path discovery when the cloud registry is not part of the flow. - Fixed direct root installs from GitHub so a missingname:in root-levelSKILL.mdfalls back to the actual repository name instead of the temporary clone directory name. - Restored the expected indentation inselftune registry --helpso the usage block matches the rest of the CLI help formatting. - Polished the cloud GitHub settings experience with branded action buttons, clearer installation action states, a consolidated production setup runbook, and lowercaseselftunebranding on key cloud surfaces. - Added signed GitHub webhook intake plus registry source metadata fields so GitHub-origin publishes can be tracked separately from CLI-pushed versions. - Hardened GitHub webhook handling so tag patterns reject unsafe multi-wildcard shapes and webhook deliveries return immediately while publish processing continues asynchronously.
- Moved canonical eval sets, generated unit tests, and unit-test run results
into SQLite as the primary local source of truth for creator-loop readiness. -
Kept mirroring those artifacts into the legacy
~/.selftune/eval-sets/and~/.selftune/unit-tests/JSON files so existing file-based workflows and commands still work during the transition. - Updated readiness/status surfaces to prefer SQLite-backed artifacts instead of depending on filesystem existence checks.
- Updated dashboard-triggered
generate-evalsto pass the canonical~/.selftune/eval-sets/<skill>.jsonoutput path explicitly instead of relying on a relative fallback filename. - Updated dashboard-triggered
generate-unit-teststo pass the canonical~/.selftune/unit-tests/<skill>.jsonpath explicitly as well, keeping readiness artifacts out of the repo working directory.
- Fixed local dashboard rollback actions to spawn
selftune evolve rollbackwith the expected proposal arguments, matching the actual CLI command surface. - Added a dashboard regression test that asserts the rollback action uses the
evolve rollbacksubcommand shape.
- Removed the forced background fill from the sticky
Evolutionheading in the shared skill report evidence rail so proposal views keep the intended transparent panel treatment while scrolling.
- Added a shared dashboard action instrumentation layer so creator-loop
commands can emit structured step progress, LLM call progress, and
provider-normalized runtime metadata without hard-coding the dashboard to one
provider. - Wired
selftune eval generateandselftune eval unit-test --generateinto that shared observer path so the live-run screen can show load/build/write steps plus provider/model/duration updates instead of only terminal output. - Generalized the live-run UI from replay-only wording to a broader action-progress surface while keeping replay as the richest source of token and cost detail.
- Added cached update availability metadata to the local dashboard health
surface so the dashboard can tell the difference between up-to-date,
auto-update-capable installs and manual-refresh source-tree installs. - Added
a passive
Update availablestatus chip in the local dashboard footer plus a dedicated update panel on/status, keeping version visibility available without polluting live creator-loop transcripts.
- Fixed proposal selection so opening a proposal link no longer gets overwritten by an automatic fallback selection. - Removed eager proposal auto-focus during initial load to keep deep links stable. - Kept readiness-driven action prioritization aligned with the active proposal focus state so child action sections no longer shift unexpectedly.
- Suppressed unsupported auto-update chatter during local source-tree runs so
dashboard-triggered creator-loop actions no longer flood the live log with
manual refresh instructions. - Updated OpenCode ingest to support the current
SQLite schema, including
time_createdtimestamps and JSON-backed message rows, instead of assuming legacycreated/contentcolumns.
- Added a live action feed in the local dashboard so creator-loop runs show
start, progress, and finish states instead of only appearing after the next
data refresh. - Added a dedicated live-run screen for creator-loop actions so
replay dry-runs can stream output, show parsed lift summaries, and display
model/platform/token context beside the terminal log. - Added structured
replay metrics to the live dashboard stream so Claude runtime replay now
reports per-run platform, model, token, cost, and duration data in real time
instead of only terminal text. - Added per-eval replay progress streaming and
SSE backfill so the live-run screen can show
eval n/N, query snippets, and pass/fail evidence even when you open the page after the run has already started. - Added dashboard action buttons for the main creator loop on skill reports: generate evals, generate unit tests, replay dry-run, baseline measurement, deploy, and watch. - Added a shared local action stream so supported terminal-runselftunecommands also appear in the dashboard without being launched from the UI. - Fixed replay dry-runs so validatedevolve --dry-runruns surface as success in the live dashboard feed even when the CLI exits non-zero to avoid accidental deployment.
- Repaired the OSS publish pipeline so npm releases can still generate SBOMs, GitHub tags, and enriched release notes even when a publish partially succeeds. - Blocked cloud dashboard indexing and added changelog coverage enforcement so shipped product changes are documented before they merge. - Opened registry publishing and rollback to Pro plans so solo skill creators can publish and iterate without upgrading to Team first. - Tightened the local dashboard skill report around proposal deep links, kept proposal-focused layouts stable while report data loads, prevented raw ENOENT errors during SPA reloads, and restored full-width creator loop layout on overview. - Unified cloud and OSS skill report styling around the shared trust status language by restoring trust panel order, removing leftover success-green treatments, and switching trust badges to the app-wide dot-and-pill status treatment.
- Added universal hook adapters for Codex, OpenCode, and Cline so selftune can capture real-time telemetry beyond Claude Code. - Added cold-start suspicion and Claude runtime replay validation to make trigger diagnostics more trustworthy when a skill has little history. - Hardened OpenCode installation so hook setup follows current plugin and config behavior instead of relying on rejected config keys. - See the OSS releases for package artifacts and per-version compare links.
- Overhauled dashboard, trust, and creator-facing contribution surfaces so health signals are easier to interpret during active iteration. - Tightened the autonomous evolve and audit path to close reliability gaps in proposal rollout and monitoring. - Added CLI auto-update, richer structured errors, description quality scoring, and unblock suggestions for faster operator recovery.
- Added full skill body evolution so selftune can refine routing tables and larger skill bodies instead of only short descriptions. - Added synthetic eval generation to help new skills bootstrap without waiting for a large session history. - Introduced cheaper validation loops, activation rules, specialized agents, and a live local dashboard server for faster iteration. - Read more in the evolution concept guide and the dashboard command reference.
- Added
selftune statusandselftune lastso you can check skill health without opening the full dashboard. - Added a local dashboard and Claude transcript backfill to make retroactive analysis practical on existing projects. - Added opt-in community export so you can share anonymized signals back to the ecosystem.
- Shipped the initial CLI with
init,grade,eval,evolve,watch,doctor, and platform ingest commands. - Added Claude Code hooks for prompt capture, skill evaluation, and end-of-session telemetry. - Introduced the initial observe → detect → evolve → watch loop that the rest of the product builds on today.