baikai-kit-0.2.0.0: CHANGELOG.md
# Changelog
All notable changes to baikai are recorded here.
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
## [baikai 0.6.0.0] - 2026-08-28
### Added
- `baikai`: `Baikai.ThinkingLevel.parseThinkingLevel :: Text -> Maybe
ThinkingLevel` and `Baikai.Evidence.parseEvidenceStrength :: Text -> Maybe
EvidenceStrength`, each beside its renderer. Three hand-copied tables — the
evidence schema's level parser, `baikai-agent`'s KDL `effort` decoder, and its
`--require-evidence` parser — now read them instead, so a level or strength
added later cannot be added in one place and missed in three. (REV-2 G.6.)
- `baikai`: `Baikai.Agent.AgentRunResult` exports its selectors (`provider`,
`exitCode`, `stdout`, `stderr`, `duration`). It exported neither them nor its
constructor, so a consumer without generic-lens could not read a run's exit
code at all. (REV-2 G.6.)
- `baikai`: `Baikai.Api.normaliseApi :: Api -> Api`, which collapses a `Custom`
tag that spells a built-in API onto that constructor. The registry applies it
to the key it stores and to the tag it is asked for, so a handler registered
under `Custom "anthropic-messages"` answers a model tagged `AnthropicMessages`
and the reverse; the two used to be separate entries and dispatch depended on
which spelling the model happened to carry. Derived `Eq`/`Ord` on `Api` are
deliberately unchanged: altering them would silently rearrange every
`Map Api` a consumer holds. (REV-2 G.4.)
- `baikai`: `Baikai.Header`, a new module exporting `HeaderName` with
`headerName` and `renderHeaderName`. See the `headers` retype under Changed.
- `baikai`: `Baikai.Error.ErrorCategory` gains `ContentFiltered` (wire tag
`content_filtered`, never retryable) with the smart constructor
`contentFiltered`. OpenAI's `finish_reason: "content_filter"` and Anthropic's
`refusal` stop now carry it. Both used to be `OtherError`, so the only way to
tell a filtered response from any other non-retryable failure was to match on
the message text. __Breaking__ for a consumer whose `case` over
`ErrorCategory` is exhaustive without a wildcard. (REV-1 1.7 residual.)
- `baikai` (breaking to construct, not to read): every record that can still
grow a field is now built from an exported base value and refined by record
update, and its constructor is no longer exported —
`Baikai.Provider.Registry.ApiProvider` (`apiProvider` /`apiProviderWith`),
`Baikai.Evidence.ModelCallEvidence` (`baseEvidence`),
`Baikai.Evidence.EvidenceRequest` (`evidenceRequest`), `Baikai.Tool.Tool`
(`mkTool`, with `emptyTool` kept for fixtures),
`Baikai.Embedding.EmbeddingModel` (`emptyEmbeddingModel`),
`Baikai.Cost.Log.CallLogConfig` (`callLogConfig`),
`baikai-trace-otel`'s `OtelSinkOptions` (`defaultOtelSinkOptions`), and
`baikai-agent`'s `AgentCliOptions` (`agentCliOptions`), `AgentCliRun`
(`agentCliRun`), `AgentJob` (`agentJob`) and `AgentConfigPaths`
(`emptyAgentConfigPaths`). Selectors, record update, `OverloadedRecordDot`
reads and generic-lens labels all keep working; only construction from the
constructor stops. Adding `describeThinking` to `ApiProvider` in 0.5.0.0 broke
every third-party registration site, and `strengthCeiling` would have broken
them again; from this release such an addition is a minor bump. (REV-2 G.1.)
- `baikai`: `Baikai.Provider.apiProvider`, which builds an `ApiProvider` from an
`Api` tag and a streaming producer, deriving `complete` with
`streamingComplete`; and `Baikai.Provider.Registry.apiProviderWith`, which
takes the completer explicitly. Both default `describeThinking` to
"nothing requested, nothing translated" and `strengthCeiling` to
`EvidenceRequestedOnly`, matching `declaredStrength (Custom _)`.
- `baikai`: `Baikai.Tool.mkTool` — a tool from its name, description and JSON
Schema. A tool built from `emptyTool` and sent unchanged reaches the wire with
`input_schema: null`; `mkTool` has no such shape.
- `baikai`: `Baikai.Agent.AgentOutputFormat` (`TextFormat`, `JsonFormat`) with
`renderAgentOutputFormat` and `parseAgentOutputFormat`, and
`AgentRunRequest.outputFormat`, defaulting to `TextFormat`. `baikai-claude`
renders `--output-format json` and `baikai-openai` renders `--json`, both
right after the effort flags; `baikai-agent` reads it from
`jobs.<name>.output-format`. This is the one setting an evidence record needs
in order to observe a run's session, model and usage, and asking for it used
to require the `provider-args` channel that an operator ceiling closes by
default — an operator should not have to open a privileged channel to get a
record. (REV-2 F.14.)
- `baikai`: `Baikai.Agent.AgentCeiling` gains three fields and the module gains
the vocabulary they need. `allowedTools :: [Text]` names tool grants the
operator permits beyond the ones `toolGrantsImpliedBy` (also new) says a
capability implies on its own; `maxTimeout :: Maybe NominalDiffTime` and
`maxOutputLimit :: Maybe Int` bound what any job may request, the second
defaulting to the new `defaultMaxOutputLimit` (67108864, sixty-four
mebibytes). `Baikai.Agent.ceilingViolations` is `applyAgentCeiling`'s violation
list on its own, so a caller can concatenate it with violations of its own.
(REV-2 F.3.)
- `baikai`: `Baikai.Content.toolArgumentsFromText` and
`Baikai.Content.isCutOffToolCall`. The first is the single rule that turns a
tool call's accumulated argument text into its `arguments` value — empty text
is an empty object, non-empty text that does not decode is kept verbatim as a
`String` — and both provider assemblers and core's stream-recovery path now
use it, so the second means the same thing at every layer.
- `baikai`: new exposed module `Baikai.Provider.Internal.StreamWorker` — the
bounded hand-off both HTTP providers now use between their SSE worker thread
and the consumer draining the stream. `FrameQueue` is a 64-slot `TBQueue` plus
a closed flag; `forkFrameWorker` closes the queue however the body ends, and
`withFrameWorker` runs the consumer under `Stream.bracketIO` so the worker is
killed when the stream stops. The module is exposed like
`Baikai.Provider.Cli.Internal`, outside the PVP promise. See
[docs/adr/0010](docs/adr/0010-a-stream-consumer-that-stops-owns-cancelling-the-producer.md).
- `baikai`: every Anthropic model in the generated catalog now carries an
explicit `CompatAnthropicMessages` record stating the two request-shaping
facts of its generation: `AnthropicMessagesCompat.thinkingStyle` (which
extended-thinking wire shape it accepts) and the new
`AnthropicMessagesCompat.supportsSamplingParameters` (whether it accepts
`temperature`, `top_p` and `top_k`). Both are sourced from
`baikai/data/models/anthropic.json`, which the fetcher writes from its
curated `anthropicInclude` table, and `baikai-gen-models` now refuses an
`anthropic-messages` entry that reaches it without a `compat` block rather
than falling back to host auto-detection, which cannot know a generation.
This is what fixes `claude-sonnet-5`, whose thinking requests were shaped by
a prefix table that did not know the id. See
[docs/adr/0009](docs/adr/0009-provider-capability-facts-live-in-the-generated-catalog-record.md).
- `baikai`: two new `Baikai.Evidence.ThinkingAdjustment` constructors,
`SamplingDroppedUnsupportedModel` and `SamplingDroppedUnsupportedApi`, encoding as
`{"kind":"sampling_dropped_unsupported_model","fields":["temperature","top_p"]}` and
`{"kind":"sampling_dropped_unsupported_api","fields":["seed"]}`. They record sampling
parameters removed because the model generation rejects them, or because the API has no
such field on any generation. Both carry a `fields` array and no `requested` level, so
they can appear on a call whose thinking mode is `absent`.
- `baikai`: `Baikai.Evidence.weakensThinking`, which says whether an adjustment weakens the
thinking the caller asked for. Strict evidence mode filters through it, so a dropped
sampling parameter is recorded without refusing the call — the documented contract is
refusing a call that would weaken the requested *thinking level*.
- `baikai`: new exposed module `Baikai.Url` — the one place baikai turns a URL
into a host name. `parseUrl` yields a `UrlParts` record with the scheme, host,
port and path, plus flags saying whether userinfo, a query string or a
fragment were present; it never holds their text, so the value cannot carry a
secret into a log line. Alongside it: `urlHost`, `hostMatchesSuffix` (moved
from `Baikai.Compat`, which now re-exports both), `renderEndpoint`,
`stripApiVersion`, and `baseUrlProblem`, which says why a URL is unusable as a
`Model.baseUrl` and what to do instead. See
[docs/adr/0008](docs/adr/0008-one-url-host-parser-and-every-consumer-uses-it.md).
- `baikai`: new exposed module `Baikai.Provider.Transport.Classify` — the one
rule every HTTP provider uses to classify a transport failure, exporting
`classifyTransportException` plus the per-type functions it composes. The rule
is *where* the failure happened, not what type it is: anything that breaks or
ends the connection after the request went out is `TransientError`, anything
that says the request or the configuration is wrong is not retryable, and a
programming error stays `OtherError`. It understands all three shapes
`http-client` can deliver — an `HttpException` of any constructor, a raw socket
`IOException`, and a raw or wrapped `TLSException` — because the manager wraps
the connect phase but not the body reader. Core gains direct `build-depends` on
`http-types` and `tls`, both already in its install plan. Written for
third-party `Custom` providers built on `http-client` as much as for baikai's
own two. See
[docs/adr/0011](docs/adr/0011-core-owns-transport-failure-classification.md).
- `baikai`: `Baikai.Error.parseHttpDate` and `Baikai.Error.retryAfterSecondsAt`.
The first parses an HTTP-date in the IMF-fixdate form servers must send plus
the two obsolete forms a recipient must accept; the second converts a
`Retry-After` header in either of its forms to seconds against a reference
instant, clamping a date already in the past to `0`.
`parseRetryAfterSeconds` keeps its integer-only contract, now a deliberate
division of labour rather than a limitation.
- `baikai`: new exposed module `Baikai.Http` — `canonicalBaseUrl`,
`getClientEnvCached` and `cachedClientEnvCount`, the process-global
`ClientEnv` cache that both HTTP provider packages now share instead of each
keeping its own. Core gains direct `build-depends` on `servant-client`,
`http-client` and `http-client-tls`, which were already in its install plan
through the `openai` SDK.
- `baikai`: `Baikai.Evidence.ThinkingModeNotTranslated`, encoded as
`"not_translated"`, and `Baikai.Evidence.untranslatedThinking`; and
`Baikai.Evidence.Build.requestedTranslation`. A path where no adapter ran to
translate the caller's level now records the level and says the translation is
unknown, instead of saying nothing was asked. (REV-2 D.2.)
- `baikai`: `Baikai.Evidence.Build.missingEvidenceError`,
`Baikai.Evidence.Build.strictnessOf` (moved here from `Baikai.Trace`, where it
was private), `Baikai.Stream.requireEvidenceOnTerminal` and
`Baikai.Provider.Registry.requireEvidenceOnResponse`. (REV-2 D.3.)
- `baikai`: `Baikai.Evidence.usageEnvelope`, and
`Baikai.Evidence.Build.endpointIdentityAt`, `prepareEvidenceAt` and
`minimalEvidenceAt`, which take the base URL the adapter actually resolved.
The three unsuffixed functions remain and pass the model's own field.
(REV-2 D.8, D.11.)
- `baikai`: `Baikai.Evidence.deriveStrength`, the single rule that turns an
observed model, a provider request id and a response id into an
`EvidenceStrength`. (REV-2 D.10.)
### Changed
- `baikai`: catalog refresh. `claude-opus-5` joins the curated Anthropic include
set (adaptive thinking, sampling parameters rejected — the facts
`docs/plans/60-make-anthropic-thinking-style-and-sampling-support-catalog-driven.md`
said whoever curated it in would have to state), and the `gpt-5.6` family
picks up its price cut: `gpt-5.6` and `gpt-5.6-sol` to $4.00/$20.00,
`gpt-5.6-terra` to $2.00/$12.00, `gpt-5.6-luna` to $0.20/$1.20 per Mtok, cache
rates in step. `Baikai.Models.Generated` gains `anthropic_claude_opus_5` and
now carries 36 enabled models. No OpenAI id was added: the `gpt-5.6` family is
still the newest one models.dev reports that speaks
`openai-chat-completions`.
- `baikai` (breaking): `ResponseFormat`'s `JsonSchema` carries a
`JsonSchemaFormat` record — `name`, `schema`, `strict`, exported
selector-only with the base `jsonSchemaFormat name schema` — instead of
holding the three fields directly. As fields of a sum they were partial
selectors: `name f` on a `JsonObject` crashed at runtime rather than failing to
typecheck, which contradicted the module's own documentation.
`-Wno-partial-fields` is dropped from the module. The JSON encoding is
deliberately unchanged (`{"tag":"JsonSchema","name":…,"schema":…,"strict":…}`)
and is now pinned by a test, because `Options` derives `ToJSON` through it and
at least one consumer keys a cache on the result. (REV-2 G.2.)
- `baikai`: `Baikai.Context.appendToolResult` returns its input context
unchanged, and runs no dispatcher, when the response is error-shaped. A failed
call has no assistant turn worth replaying and no tool calls to answer;
appending its empty message put a turn into the transcript the model never
took. `runToolLoop` has always stopped on such a response — the documented
direct round trip in `docs/user/tools.md` reaches `appendToolResult` instead,
and now behaves the same way. Its Haddock also stops claiming multi-call
concurrency lives in the dispatcher: the calls are traversed in order.
(REV-2 G.7.)
- Release metadata (REV-2 G.8): every publishable package now declares
`tested-with: GHC ==9.12.4` and ships its `CHANGELOG.md` (a symlink to the
root one, as `baikai` already did) via `extra-doc-files`, so Hackage shows a
changelog and a tested compiler for all seven. `baikai-claude` and
`baikai-openai` describe what they actually contain — four surfaces each, not
"wraps package X" — and `baikai-trace-otel`'s `streamly-core` bound is
`>=0.3 && <0.5`, matching every other package in the workspace rather than
excluding the 0.4 series the others accept.
- `baikai` (breaking): `Options.headers` and `Model.headers` are keyed on
`Baikai.Header.HeaderName` — a newtype over a case-insensitive `CI Text` that
keeps the original spelling — instead of `Text`. A header name is
case-insensitive on the wire, so a `Map Text Text` holding both
`Authorization` and `authorization` sent whichever the assembling fold reached
last; the map now holds one entry per header and the last write wins, as a
caller writing two spellings would expect. `HeaderName` has an `IsString`
instance, so `Map.singleton "x-test" "1"` and `#headers` updates keep
compiling; the spelling given is what goes out on the wire and into JSON.
(REV-2 G.5.)
- `baikai` (breaking): `Options.stopSequences` is `[Text]`, where empty means
"send nothing", instead of `Maybe (Vector Text)` — `Nothing` and `Just []`
were indistinguishable on the wire and only one of them could be right. Plan
43's rule is lists for caller-side configuration and `Vector` for
provider-bound sequences; this was the one field breaking it. `Options.seed`
is `Maybe Int` rather than `Maybe Integer`: a seed is a machine integer at
every provider that accepts one, and it now sits beside
`timeoutMs :: Maybe Int`. (REV-2 G.5, R14.)
- `baikai` (breaking): `StopReason.Aborted` is removed. Nothing produced it —
timeouts are `ErrorReason`/`TransientError`, and a consumer abort is recorded
as evidence `CallAborted` — while `responseError`, `eventsFor` and
`runToolLoop` all treated it as a *success*, so a value that reached any of
them would have been silently mishandled. Since 0.6.0.0 a stream consumer that
stops cancels the producer, so no consumer is left to receive such a terminal
either. (REV-2 B.6.)
- `baikai`: dispatching a model whose `api` is still `emptyModel`'s
`Custom ""` says so — `No provider registered for API: <blank Custom tag —
emptyModel.api was never set>` — where the message used to end after the
colon. `emptyModel`'s Haddock says the same thing. (REV-2 G.4.)
- `baikai`: `withTrace` and `withTraceStream` wait at most one second for the
trace sink after writing the shutdown sentinel. On expiry the worker is
abandoned — not killed, which would abort the sink's fold mid-step and lose
its end-of-stream action — the call proceeds, and one stderr line reports
`the trace sink did not confirm delivery within 1000 ms; its worker was
abandoned, and events already queued may still be delivered later`. A sink
that blocked forever used to hold the call forever and swallow the first
attempt to cancel it. A caller under `EvidenceRequired` whose sink did not
confirm delivery gets a failed call, through the same path a throwing sink
takes; `Baikai.Evidence.Build.sinkFailureError` now says "its record was not
confirmed written" rather than "not written", which is the honest claim for
an abandoned worker whose events are still queued. The synthetic terminal a
consumer's abort produces is delivered from a garbage-collection hook and is
not guaranteed before process exit; that was always true and is now stated in
`docs/user/model-call-evidence.md`, `docs/capabilities/call-tracing.md` and
the `Baikai.Trace` module documentation, with the pattern for callers who need
the record. See
[docs/adr/0015](docs/adr/0015-trace-cleanup-is-bounded-and-abort-cleanup-is-gc-eventual.md).
(REV-2 D.5, Theme 7.3.)
- `baikai`: `Baikai.Trace.Sink.multiSink` runs each member on its own drain
thread behind its own unbounded channel, instead of folding `Fold.tee` across
the list. `Fold.tee` runs one member then the other and lets either's
exception escape, so a single throwing member stopped delivery to every
sibling for the rest of the call and skipped their end-of-stream actions — an
OpenTelemetry span paired with an unwritable file sink was opened and never
ended, and nothing was exported. The step never blocks; the final action sends
every member the sentinel, waits for every member, and reports one aggregate
failure naming each failed member by zero-based index
(`1 of 2 member sinks failed: member 0: …`). (REV-2 D.6.)
- `baikai`: `AgentSafety.allowedTools` is documented as the __grant__ it is.
On Claude Code it renders `--allowedTools`, whose help reads "list of tool
names to allow": it pre-approves tools the permission mode would otherwise
raise a request for, and in an unattended run a request nobody answers is
denied. The old Haddock called it "optional narrowing of the provider's tool
set", which was the opposite, and `applyAgentCeiling` never looked at it. It
is now bounded: a grant passes when the maximum capability implies it
(`read-only` implies `Read`, `Glob`, `Grep`, `NotebookRead`, `TodoWrite`;
`edit-workspace` adds `Edit`, `MultiEdit`, `Write`, `NotebookEdit`;
`full-access` implies every grant) or when the operator named it in
`policy.allowed-tools`. Matching is exact, so `Bash(git *)` is not `Bash`.
A repository job that grants itself `Bash` under `edit-workspace` — which
passed unexamined before — is now refused with exit 77 before any process is
created. (REV-2 F.3.)
- `baikai` (breaking): `Baikai.Agent.CeilingViolation` gains five constructors:
`ToolGrantForbidden`, `TimeoutExceeded`, `OutputLimitExceeded`,
`RepositoryScopeForbidden` and `WorkingDirOutsideRepository`. A `case` over
the type that was exhaustive is no longer.
- `baikai` (behaviour): the default ceiling has a finite `maxOutputLimit`, so
`applyAgentCeiling defaultAgentCeiling` now refuses a request whose
`outputLimit` is `Nothing` — capture without bound is exactly what the
maximum exists to refuse. Jobs resolved through `baikai-agent` are unaffected:
that layer's own default supplies a finite limit, and only an explicit
`output-limit "unlimited"` reaches the ceiling as `Nothing`.
- `baikai`: a tool call cut off by the output cap is no longer executed.
`runToolLoop` stops with the response and its tool calls intact when any call
is cut off, and `appendToolResult` appends a `ToolResultMessage` with
`isError = True` explaining why instead of calling the dispatcher. Previously
both assemblers replaced truncated arguments with `{}` and a tool loop
happily ran the call with no arguments at all. (REV-2 B.2.)
- `baikai`: `Baikai.Model.anthropicMessagesCompatFor` no longer overlays a
thinking style guessed from the model id onto a model whose `compat` is
`CompatNone`. `CompatNone` now means host auto-detection alone — the budget
thinking shape, sampling parameters supported. Every catalog model carries an
explicit record, so this changes nothing for them; a **hand-rolled** model
naming an adaptive-era id (`claude-sonnet-5`, `claude-opus-4-7`,
`claude-opus-4-8`, `claude-fable-5`) must now carry
`CompatAnthropicMessages (defaultAnthropicMessagesCompat {thinkingStyle = AnthropicThinkingAdaptive, supportsSamplingParameters = False})`
or start from the catalog value.
- `baikai`: `Baikai.Evidence.evidenceSchemaVersion` is now
`baikai.model-call-evidence/1.1`. A minor bump: the two sampling adjustment kinds are a
compatible addition, and no previously recorded digest changes.
- `baikai`: HTTP 413 classifies as `ContextOverflow` rather than `OtherError`,
from the status alone and whatever the body says. 413 *is* the size-limit
status and the caller's remedy — shrink the input — is the same either way;
making the category depend on body wording would recreate for 413 the
inconsistency this release fixes for connection resets. (REV-2 A.7.)
- `baikai`, `baikai-claude`, `baikai-openai`: an HTTP-date `Retry-After` is
converted to seconds instead of ignored. Both transports use the response's own
`Date` header as the reference instant, falling back to the local clock, so a
CDN-fronted `429` — the common case for a date-valued `Retry-After` — now
carries a hint rather than leaving the caller to guess. (REV-2 A.9.)
- `baikai`: **breaking.** `Baikai.Embedding.EmbeddingModel.apiKey` is now
`Maybe ApiKeySource` rather than `ApiKeySource`. `Nothing` means the
conventional environment variable for the model's host, from
`defaultApiKeyEnvForBaseUrl` — the same table the chat providers use — and a
host that table does not know refuses with an `AuthError` naming
`EmbeddingModel.apiKey`. Migration: `apiKey = source` becomes
`apiKey = Just source`. `EmbeddingModel` also derives `Eq` and `Generic`, so
the `#field .~ value` idiom works on it as it does on every other record.
(REV-2 E.3.)
- `baikai`: **breaking.** `AgentRunFailure`'s `RunTimedOut` constructor now
carries a new record `AgentTimedOut` — the configured `limit` plus the
`stdout` and `stderr` a timed-out run drained before its process group was
killed — instead of a bare `NominalDiffTime`. A caller matching
`RunTimedOut limit` becomes `RunTimedOut timedOut` and reads `timedOut ^.
#limit`; `renderAgentRunFailure` is unchanged in what it says. The bytes were
always there, drained from the moment the child was spawned, and were simply
dropped on the timeout path — which is the run an operator most wants an
account of, because the tool started, may have consumed tokens, and may
already have changed the working tree.
- `baikai`: under `EvidenceRequired`, a successful terminal that carries no
evidence record fails the call with `missingEvidenceError` rather than
returning a silent success with zero `call_evidence` lines. Strict mode
guaranteed that a record which was built and then lost fails the call; it did
not guarantee that one was built. The rule is applied at both dispatch points,
so `completeRequest` with no sink gets the same guarantee as a streaming call;
a failed call keeps the provider's own error, and best effort is unchanged.
See `docs/adr/0014-strict-evidence-means-a-record-exists.md`. (REV-2 D.3.)
- `baikai`: a caller's thinking level is recorded on every evidence path — the
consumer abort, an unregistered provider, a `complete` handler that threw, and
each provider's `immediateError`. The abort path asks the registered adapter's
own `describeThinking`; the others record `not_translated`. All four used to
record the caller's request as `absent`, which
`docs/adr/0002-requested-translated-observed-are-never-collapsed.md` forbids.
(REV-2 D.2.)
- **`baikai.model-call-evidence/2.0`.** Two digests cover different bytes, so a
verifier must now select its rules by `schema_version`. `response_commitment`
covers the provider-reported token counts and never baikai's computed cost:
the cost comes from the caller's catalog rates rather than from the response,
so the digest used to change whenever a price was edited and a verifier
holding only the response could not recompute it. `request_configuration`
summarises `output_config` and `response_format` as it already summarised
`tools`, because a structured-output JSON schema carries author-written
`description` strings and is content wherever it appears — the same schema was
stripped from `tools[].input_schema` and survived verbatim through the other
two keys. `thinking.mode` may also now be `"not_translated"`, which is a
compatible addition. (REV-2 D.7, D.11.)
- **Breaking.** `baikai`: `Baikai.Provider.Registry.ApiProvider` gains a fifth
field, `strengthCeiling :: EvidenceStrength`, and
`Baikai.Evidence.Build.checkEvidenceRequirements` takes that ceiling where it
took an `Api`. The gate compared against `declaredStrength`, a table keyed by
the API tag, which necessarily answered `EvidenceRequestedOnly` for every
`Custom` transport — so a gateway that genuinely observes a model could never
satisfy a strict caller who required that it did. Only a provider knows what
its evidence reaches. `EvidenceRequestedOnly` reproduces the old behaviour for
any custom provider; the four built-in providers fill the field from
`declaredStrength`, which is unchanged in value and still used by the
unattended-agent surface. (REV-2 D.10, G.1.)
- `baikai`, `baikai-claude`, `baikai-openai`: one strength derivation replaces
three. An observed **response id** now counts as correlation alongside a
captured request-id header, so a host that names its model and its response id
on every chunk but sends no header reaches `model_observed` instead of
`requested_only` — which had put it *below* a host that sent only a header and
named nothing. `anthropicStrength` and `openaiStrength` are removed;
`Baikai.Provider.Cli.Internal.subprocessStrength` keeps its signature and
delegates. (REV-2 D.10.)
### Removed
- `baikai` **0.6.0.0** (breaking): the sixteen `_Type` base-value aliases deprecated in
0.3.0.0 — `_Options`, `_Context`, `_Model`, `_ModelCost`, `_Response`,
`_Usage`, `_Cost`, `_CostBreakdown`, `_Tool`, `_TextContent`,
`_ThinkingContent`, `_ToolCall`, `_ImageContent`, `_EmbeddingModel`,
`_InteractiveLaunchRequest` and `_InteractiveLaunchResult`. Each has an
`empty…` or `zero…` replacement of the same value, named in the pragma that
has been on it since 0.3.0.0. The 0.3.0.0 entry said they remained "for this
release"; 0.4.0.0 and 0.5.0.0 shipped without removing them because no entry
named a version.
`docs/adr/0016-deprecated-names-are-removed-at-the-next-major.md` now fixes
the rule: a name deprecated in `A.B.0.0` is removed in `A.(B+1).0.0`, and
every pragma says so. (REV-2 G.3.)
- `baikai` **0.6.0.0** (breaking): `Baikai.Trace.newEventId`. It has delegated to
`Baikai.Evidence.newCallId` since 0.5.0.0; call that. (REV-2 G.3.)
- `baikai` **0.6.0.0** (breaking): `Baikai.Compat.defaultAnthropicThinkingStyle`, deprecated
earlier in this cycle. Nothing in baikai consults it — the thinking style of a
first-party Anthropic model is a field of its generated catalog record
(`Baikai.Models.Generated`); start from that value, or set
`CompatAnthropicMessages` explicitly.
- `baikai` (breaking): `AgentRunRequest.envPassthrough` is renamed `envRequires`.
The field is a list of variables the job declares it requires, checked as a
precondition; it has never passed anything through, and the KDL key has said
`env-requires` since the setting existed.
- `baikai` (breaking): `AgentRunFailure.OutputMalformed`, and with it
`baikai-agent`'s exit code 70 and its `internalExitCode` export. Nothing ever
constructed the constructor, and giving it a producer would have been wrong:
the runner treats the tool's output as best-effort observation and its
deliverable is the changed working tree, so a run that edited files correctly
and then printed an unparseable final line would have been reported as a
failure with its exit code and output discarded. A record's `strength` and
`unobserved` fields already say when output could not be read. (REV-2 F.13.)
### Fixed
- `baikai`: the terminal event and its evidence record are pushed to the trace
sink exactly once under asynchronous exceptions. The terminal path pushed the
evidence record, pushed the terminal event and only then set the
already-sent flag; an exception delivered between the last two made the
stream finaliser read the flag as unset and push a second `CallEvidence` and
an `aborted` `CallFailed` after the real `CallFinished`, so a sink saw two
records and two contradictory terminals for one call. All three writes now
run inside one `uninterruptibleMask_` with the flag first. (REV-2 D.4.)
- `baikai`: `Baikai.Cost.Log.closeCallLog` is idempotent. The first caller
claims the handle and waits for the worker; a second returns at once instead
of blocking forever on an `MVar` the worker had already emptied — a shape
`withCallLog` makes easy to reach, since its bracket closes a handle the body
may also have closed. An `appendEntry` after the close enqueues nothing.
- `baikai`: `reassembleResponse` is total under duplicated, late and
timestamp-less input. The first `EventStart` wins the skeleton and
`responseId` merges with `<|>`, so a later `Nothing` cannot erase an id an
earlier event supplied; events after the first terminal are ignored, so a
producer that keeps talking cannot rewrite the answer; and `latencyMs` falls
back to the reassembler's own wall clock when neither the skeleton nor the
terminal carries a provider timestamp, instead of reporting a zero that reads
as "instant". (REV-2 B.7.)
- `baikai`: an `EmbeddingModel` pointed at a non-OpenAI host no longer sends
`OPENAI_API_KEY` to it. The default key source was that variable whatever the
base URL said, so pointing the client at DeepSeek handed DeepSeek an OpenAI
credential. It now resolves per host, and refuses an unknown one. New
`resolveEmbeddingKey` and `embeddingClientEnv` expose both decisions without
making a request. (REV-2 E.3.)
- `baikai`: `Baikai.Embedding.embed` no longer allocates a TLS manager per call.
It used the `openai` SDK's own `getClientEnv`, which builds a fresh manager
every time; it now takes one from `Baikai.Http`'s process-global cache, the
same one the chat providers use, so an embedding call and a chat call to one
host share a connection pool.
- `baikai`: **a credential in a header is no longer printed.** `Options.headers`
and `Model.headers` went through derived `Show` and `ToJSON` instances that
rendered every value verbatim — while `Baikai.Options`' own documentation
invites callers to put a gateway's `Authorization` header there and the
getting-started guide tells them to `print resp`, which renders the embedded
`Model`. Both types now have hand-written instances that render exactly what
the derived ones did, except that the value of a header whose name looks
credential-carrying (`authorization`, `api-key`, `apikey`, `token`, `secret`,
`cookie`, `password`, or any name ending in `-key`, case-insensitively) prints
as `<redacted>`. `Baikai.Auth` exports the three pieces — `redactedMarker`,
`isCredentialHeader`, `redactHeaderValues` — so a caller can apply the same
rule to its own logging. Only the rendering changes: the field is untouched,
`Eq` is untouched, and the header is still sent as written. A JSON round trip
of a `Model` is deliberately lossy, since a serialised `Model` is exactly the
thing that should not carry a key. (REV-2 E.2.)
- `baikai`: an API-key environment variable set to the empty string, or to
nothing but whitespace, now counts as **unset**. `ApiKeyEnv` fails with an
`AuthError` naming the variable and saying it is not set or is empty;
`ApiKeyEnvChain` skips it and continues, and reports every name when none
yields a key. Previously an empty variable resolved to an empty key, which
short-circuited a chain and produced `Authorization: Bearer ` and a provider
401 that said nothing about the cause. A key with real content is still passed
through untrimmed. (REV-2 E.6.)
- `baikai`: **the host parse no longer lets a base URL choose which key baikai
sends.** `urlHost` took the text after the *last* `@` anywhere in a URL, so
`https://proxy.example.com/v1?u=@api.openai.com` named the host
`api.openai.com`: `defaultApiKeyEnvForBaseUrl` resolved `OPENAI_API_KEY`,
`autoDetectOpenAICompletions` returned OpenAI's own compatibility record, and
the bearer token went to `proxy.example.com`. Anyone who could set `baseUrl` —
a `Model` decoded from JSON, a proxy override — could pick which provider's
credential to be handed. The same defect broke the benign direction:
`https://api.openai.com/v1/@x` named the host `x` and resolved no key at all.
The authority now ends at the first `/`, `?` or `#`, and userinfo is only ever
the last `@` inside it. (REV-2 A.1 / E.1.)
- `baikai`: `Baikai.Evidence.Build.sanitizeEndpoint` was a second, separately
written parser that bounded the authority at the first `/` only, so a URL with
a query and no path recorded the wrong host. It is now `renderEndpoint <$>
parseUrl`, which also means a recorded endpoint has a lower-cased scheme and
host; the path keeps its case and trailing slash.
- `baikai`: `parseCodexJsonlStream` assembles lines in **linear time**. It
previously unpacked every chunk into a stream of bytes and appended them one
at a time with `BS.snoc`, copying the whole accumulator per byte — quadratic
in line length, so one codex event carrying a two-million-character message
cost on the order of a trillion byte moves and in practice never finished.
Lines are now cut out of each chunk with `BS.elemIndex` and `BS.splitAt`, and
the pieces of a line that spans a chunk boundary are joined once. Behaviour is
unchanged: a non-JSON line is still skipped, and a last line without a
trailing newline is still parsed.
- `baikai`: a Codex custom agent's instructions body renders as a TOML
**literal** multi-line string (`'''`), which interprets nothing, instead of a
basic one (`"""`), which interprets backslash escapes. As a basic string an
instruction as ordinary as "match `\d+`" made Codex refuse to load the file;
`tomllib` rejects the old output with `Unescaped '\' in a string`. A body a
literal string cannot hold — one containing three apostrophes, a bare carriage
return, or a control character other than tab and newline — falls back to a
fully escaped basic string. `tomlString`, which renders `name` and
`description`, now escapes every control character as TOML 1.0 requires
instead of only the five it happened to name.
- Documentation: `baikai`'s Haddock no longer describes behaviour the code left
behind. The trace event's token counts are `Maybe` because a non-assistant
terminal has no usage, not because the CLI providers report nothing — since
0.5.0.0 both carry what the tool reported. `EventStart`'s `partial` is a
message skeleton with empty content, zero usage and no stop reason; the api,
provider and model id live on the `Response`. A lifted stream's `EventStart`
carries the final usage and stop reason already filled in, because the
response is complete before the stream begins. `Baikai.CacheRetention` no
longer mentions an OpenAI Responses 24-hour bucket no code emits. System
prompts are documented as living on `Context.systemPrompt` rather than on a
`Baikai.Request` module that no longer exists, `emptyModel`'s `compat` is
described as auto-detection rather than a placeholder, tool dispatch says
calls run one at a time in order, and every reference to a plan number is
gone. (REV-2 H.4.)
## [baikai-claude 0.6.0.0] - 2026-08-28
### Added
- `baikai-claude`: `Baikai.Provider.Claude.Internal.Request` exports `planRequest`,
`SamplingPlan`, `uncappedMaxTokensFloor` and `normalizeToolCallId` as test seams.
`planThinking` and `describeThinkingFor` are now projections of `planRequest`, so the
strict gate, the request builder and the evidence record read one answer.
### Changed
- `baikai-claude`, `baikai-openai` (breaking): each provider's streaming
machinery moved from `Baikai.Provider.<P>.Api` to
`Baikai.Provider.<P>.Internal.Stream` — the `SseDriver` seam, `liveSseDriver`,
`<p>StreamWith`, `Assembler`, `emptyAssembler`, `translate`, and on the OpenAI
side `RawChunk`, `RawToolDelta`, `parseChunk`, `parseFrame`, `TagScanState`,
`scanThinkTags`, `closeOpenStream`, `RawUsage`, `parseUsage` and
`rawUsageToUsage`. `Api` now exports exactly `register`, the provider value
and the live stream function. The `.Internal` module is exposed for the test
suites and sibling packages and, like every `.Internal` module, may change in
any release without a major bump — so changing the assembler stops being a
documented break. `Shape`, `Sse` and `Transport` keep their names and gain the
same no-guarantees header. `_TagScanState` is renamed `emptyTagScanState`.
(REV-2 G.1.)
- `baikai-claude`, `baikai-openai`: a consumer that stops reading now stops the
provider. Both packages fork their SSE worker under `Stream.bracketIO` and
hand frames through the bounded `FrameQueue` above instead of an unbounded
`Chan`. A consumer that cancels — `Ctrl-C`, `System.Timeout.timeout`,
`cancel` — releases the HTTP connection immediately; a consumer that abandons
the stream (`Stream.take 3`) stops the socket read within 64 further frames
and releases the connection at the next major garbage collection. Previously
the worker read the entire generation into memory for a consumer that would
never look at it, and the provider billed all of it. The three cleanup
strengths are stated in
[docs/adr/0010](docs/adr/0010-a-stream-consumer-that-stops-owns-cancelling-the-producer.md)
and in caller terms in `docs/user/streaming.md`.
- `baikai-claude`: `anthropic_claude_sonnet_4_6` now sends the adaptive
thinking shape rather than `budget_tokens`. The budget shape is deprecated
for that generation; baikai sends the shape Anthropic documents as current.
- `baikai-claude`, `baikai-openai`: **behaviour change.** `Options.timeoutMs` of
`Just n` with `n <= 0` is refused as `InvalidRequest` before the action runs, so
no connection is opened. `System.Timeout.timeout` returns immediately at zero
and runs unbounded below it, and the previous `max 0` clamp made both spellings
fail instantly as a *retryable* `TransientError` — a classification a caller's
retry loop re-issues forever for what is a configuration mistake. `Nothing`
remains the only spelling of "no bound". (REV-2 A.10.)
- `baikai-claude`, `baikai-openai`: an evidence record's `endpoint` names the
host the call actually went to. Both adapters substitute a vendor default for
an empty `Model.baseUrl` inside `prepareCall`, so a call with a perfectly
definite destination recorded `endpoint: null`. Where no adapter ran, `null`
remains the truthful answer. (REV-2 D.8.)
- `baikai-claude`: the `claude` dependency moves from `^>=1.4` to `^>=1.5`.
1.5.0 adds a `Pause_Turn` constructor to `Claude.V1.Messages.StopReason`, and
`mapStopReason` matches that type with no wildcard under
`-Werror=incomplete-patterns`, so the bump forced a decision. A paused turn
maps to `Stop`: Anthropic suspends the turn mid-flight for a long-running
server-side tool and expects the caller to send the message back to continue
it, so nothing failed, and `Baikai.StopReason` has no constructor that says
"resume me". Widening that public sum is a breaking change for every consumer
who matches on it exhaustively, and it is not this bump's to make. The general
rule is
[ADR 0018](docs/adr/0018-a-provider-stop-reason-with-no-baikai-equivalent-maps-to-the-nearest-truthful-one.md):
a provider stop reason with no baikai equivalent maps to the constructor that
is truthful about whether the call failed, and the sum widens only when baikai
would behave differently for it.
- `baikai-claude`: `Messages.StreamUsage` lost its `Generic` instance in `claude`
1.5.0, so the `message_delta` usage is read through `OverloadedRecordDot`
rather than a generic-lens label. `Messages.max_tokens` and
`Messages.output_config` became ambiguous selectors — `Messages.Fallback`
carries both names — so the provider's tests read them through `^. #max_tokens`
and `^. #output_config` instead.
### Removed
- `baikai-claude`, `baikai-openai` **0.6.0.0** (breaking): the eight registration shims —
`registerWith`, `registerWithRegistry` and `registerWithRegistryAndConfig` in
both `Cli` modules, and `registerWithRegistry` in both `Api` modules. Register
the exported provider value instead:
`registerApiProvider (claudeCliProvider cfg)`,
`registerApiProviderWith reg (codexCliProvider cfg)`,
`registerApiProviderWith reg claudeMessagesProvider`. The batch-mode note that
had accumulated on `registerWith` — why `complete` stays on the direct path
rather than going through `streamingComplete` — moves to the provider value it
describes. (REV-2 G.3.)
- `baikai-claude`, `baikai-openai`: `responseToError` and `classifyErrorText`
(and its private `classifySdkHttpText` half) from both
`.Internal.ErrorClass` modules. Neither package runs a `servant-client` client
on the chat path any more, so the `ClientError` branch was unreachable, and the
text classifiers parsed a string shape the local SSE transports stopped
producing in July. The phrase table `classifyErrorText` held survives as the
message fallback inside `classifyErrorFrame`, pinned through the entry point the
runtime actually uses. Both modules are documented as outside the PVP-stable
surface, so this is not a major bump; version bumps are recorded once, later.
- **Breaking.** `baikai-claude`: `Baikai.Provider.Claude.Api.anthropicStrength`
and `baikai-openai`: `Baikai.Provider.OpenAI.Api.openaiStrength`, both replaced
by `Baikai.Evidence.deriveStrength`.
### Fixed
- `baikai-claude`, `baikai-openai`: a failure that lands while the response body
is streaming is classified as the transient failure it is. A connection reset,
a server closing the socket mid-chunk, a body shorter than its declared length
and a TLS session torn down after the handshake all now terminate the stream
with `TransientError` and `isRetryable = True`, carrying whatever text had
already been drained. Every one of them used to be `OtherError` with
`isRetryable = False`, while the identical failure at connect time was
transient — because `http-client` wraps the connect phase with the manager's
exception wrapper and the body reader with nothing that converts a socket
`IOException` or a `TLSException`, so those reached the worker raw and missed
the `HttpException` branch entirely. (REV-2 A.2.)
- `baikai-claude`, `baikai-openai`: a transport failure mid-stream now closes
the blocks that were open when it arrived, on both providers, so a consumer
reading raw events and a consumer reassembling them see the same partial
output. Both providers built their terminal from the closed blocks alone and
silently dropped open text, thinking and tool arguments. On the Claude side
this covers `translate (Left …)`, the in-band `error` frame, and the
unexpected end of stream. (REV-2 B.3.)
- `baikai-claude`: an SSE frame whose event `type` — or whose
`content_block_delta` `delta.type` — the SDK has no constructor for is now
skipped instead of ending the stream with a decode error. The SDK decodes both
with no unknown-tag fallback, so a new frame type from Anthropic used to be a
terminal fault. A frame of a *known* type that still fails to decode remains
one. `Baikai.Provider.Claude.Sse` exports the new `decodeFrame`. (REV-2 B.5.)
- `baikai-claude`, `baikai-openai`: an empty `data:` heartbeat is ignored, and
on the OpenAI side `[DONE]` is compared after trailing whitespace is trimmed,
so `data: [DONE] ` and `data: [DONE]\r` end the stream rather than failing to
decode. (REV-2 A.8.)
- `baikai-claude`: every failing stream now begins with `EventStart`. The
producer pre-seeds the start event before the first wire read, exactly as the
OpenAI producer already did, and `message_start` updates the assembler without
emitting a second one. Previously a 401, a rate limit, an in-band `error`
frame or an EOF arriving before `message_start` produced a lone `EventError`,
breaking the protocol `Baikai.Stream.Event` documents. `StartPayload.responseId`
is consequently `Nothing` on both HTTP providers; the provider's message id
rides `TerminalPayload.responseId`, which `reassembleResponse` already prefers.
(REV-2 A.4, REV-1 Theme 1.1.)
- `baikai-claude`, `baikai-openai`: an asynchronous exception delivered to the
stream worker can no longer strand its consumer. End-of-frames is a flag set
by the worker fork's own `finally` rather than a sentinel value pushed onto
the channel, so a worker that dies without running its normal exit path still
ends the stream in an `EventError`. Previously the consumer blocked until the
runtime's deadlock detector noticed.
- `baikai-smoke`: two keyed cases against `claude-sonnet-5` — one asking for
thinking (which is a 400 before this release) and one setting `temperature` — plus
`deepseek-chat` and `openrouter/openai/gpt-4o-mini` in `apiCases`, so the tool and
structured-output smokes run against a compatible host that is not OpenAI.
`CompatSmoke` now asserts DeepSeek honoured the output cap rather than only that it
answered, and `CacheSmoke` asserts the cached token classes cost something.
- `baikai-claude`: a thinking request on `claude-sonnet-5` no longer 400s. It sends
`"thinking":{"type":"adaptive"}` and no `budget_tokens`, because the shape is read off
the model's catalog record rather than guessed from its id. (REV-2 C.1.)
- `baikai-claude`: `temperature` and `top_p` are no longer sent to a model generation that
rejects them with a 400. They are omitted and the omission is recorded as
`sampling_dropped_unsupported_model` in the call's evidence. `seed`, `frequencyPenalty`
and `presencePenalty`, which the Anthropic Messages API has no field for on any
generation, are recorded as `sampling_dropped_unsupported_api`. (REV-2 C.1, C.5.)
- `baikai-claude`: a model whose `maxOutputTokens` is `0` no longer sends
`"max_tokens":0`, which Anthropic rejects — and, with thinking set, no longer had its
whole thinking plan discarded for not fitting inside a ceiling of zero. It sends
`uncappedMaxTokensFloor` (1024, the SDK's own default) instead. An explicit
`maxTokens = Just 0` is still forwarded as written. (REV-2 C.2.)
- `baikai-claude`: replay no longer sends an empty text block or an empty `content` array,
both of which Anthropic rejects. An empty text block is dropped; an assistant turn left
with nothing is dropped whole (it is baikai's own artifact — a block that closed with no
deltas, or only unsigned thinking, which replay already omits); a user turn left with
nothing is refused locally with a message naming the turn. (REV-2 C.3.)
- `baikai-claude`: tool-call ids that differ only in characters the alphabet forbids, or
only past character 64, no longer normalise onto the same id and misroute a tool result.
A conforming id passes through unchanged — every id Anthropic and OpenAI actually mint
does — and any other is truncated to 51 characters and suffixed with twelve hex
characters of its SHA-256. Two `tool_use` blocks in one turn that still collide are
refused rather than sent. (REV-2 C.7.)
- Documentation: `baikai-claude`'s and `baikai-openai`'s Haddock point at the
functions that exist. `Baikai.Compat` named
`Baikai.Provider.OpenAI.Api.mkOpenAIResponseFormat`,
`…Api.applyThinkingFormat` and `…Api.translateTextLikeDelta`; the first two
moved to `…Internal.Request` and the third is
`…Internal.Stream.scanThinkTags`. `ThinkingFormat`'s note said the six
non-native shapes all clamp through `compatibleEffort`; three do, Z.ai and
Qwen send a bare toggle, and `ThinkingFormatNone` drops the control.
`immediateError` carried two `-- |` headers where one was intended.
(REV-2 H.4.)
- `baikai-claude`: an Anthropic call reports its thinking tokens. `Usage.reasoningTokens`
was hard-coded to `Nothing` on this provider because `claude` 1.4.0's
`Messages.Usage` had no breakdown to read; 1.5.0 adds
`output_tokens_details.thinking_tokens`, and both `message_start` and
`message_delta` now fill the field from it. `reasoningTokens` is an
informational subset of `outputTokens`, so no total and no cost moves.
- `baikai-claude`: the prompt-side token counts survive a server-side tool run.
The final `message_delta` used to contribute only `output_tokens`, and
`inputTokens`, `cacheReadTokens` and `cacheWriteTokens` kept whatever
`message_start` had reported — which is wrong for a call whose prompt grew
mid-stream. `claude` 1.5.0 exposes those three on `Messages.StreamUsage`, and
each is now taken when present. An absent field still keeps the
`message_start` figure rather than zeroing it, so a model that sends only
`output_tokens` is accounted for exactly as before.
## [baikai-openai 0.6.0.0] - 2026-08-28
### Added
- `baikai-openai`: `Baikai.Provider.OpenAI.Internal.ErrorClass.classifyErrorFrame`
and `Baikai.Provider.OpenAI.Api.parseFrame`, which sort a decoded SSE payload
into a classified in-band error or a completion chunk.
### Changed
- `baikai-openai`: **breaking.** `Baikai.Provider.OpenAI.Shape`'s
`injectThinkingShape`, `describeThinkingShape`, `shapeRequestBody` and
`streamRequestBody` take a `Bool` after the compat record — whether the model
advertises reasoning support (`Model.reasoning`). A level on a `reasoning = False`
model now sends no `reasoning_effort`, `reasoning`, `thinking` or `enable_thinking`
key on any host, and records `thinking_dropped_unsupported_model` instead. The model
check runs before the host-format check. This is what stops `gpt-4o-mini` plus a
level from 400ing. (REV-2 C.4.)
### Fixed
- `baikai-openai`: an in-band `{"error": …}` frame on a `2xx` stream terminates
the call with the frame's own classification, status and message. Compatible
hosts (OpenRouter, DeepSeek, Together) report an upstream failure they only
learned about after committing to a `200` this way, and `parseChunk` never
looked at `error`. The pre-fix behaviour was worse than a bad category:
OpenRouter's frame carries `choices[0].finish_reason = "error"`, which mapped
to `Stop`, so the call ended as `EventDone` with `errorInfo = Nothing` — a
consumer switching on the terminal saw a *completed* call. A frame with no
`choices` beside the error ended as
`OtherError "openai stream ended without finish_reason"`. (REV-2 A.3.)
- `baikai-openai`: reasoning that arrives after visible text closes the open
text block before opening the thinking block, so at most one of the two is
open at a time, every `_End` precedes the next `_Start`, and no `contentIndex`
is revisited after a later one. (REV-2 B.4.)
- `baikai-openai`, `baikai-claude`: **a provider POST no longer follows
redirects.** `http-client`'s default is to follow up to ten with every header
intact, so a 3xx would have re-sent the bearer token (or `x-api-key`) to
whatever host the `Location` header named. `redirectCount` is now zero and the
3xx is delivered as the one in-band terminal error carrying its status. Each
transport's request builder is exported as `buildRequest`, so the method, the
composed path and the redirect policy are assertable without a connection.
(REV-2 A.5 / E.4.)
- `baikai-openai`, `baikai-claude`, `baikai`: **the base-URL convention is
stated and enforced.** `Model.baseUrl` and `EmbeddingModel.baseUrl` are the
API *root* — the host, or the prefix a host mounts the API under — because
baikai appends `/v1/chat/completions`, `/v1/messages` or `/v1/embeddings`
itself. A trailing `/v1` is accepted and removed rather than doubled, so
`https://api.deepseek.com/v1` now requests `/v1/chat/completions` instead of
`/v1/v1/chat/completions`. A base URL with no scheme, a scheme other than
`http`/`https`, credentials, a query string, a fragment, or a path that is
already an endpoint is refused as an `InvalidRequest` naming the problem —
and refused *before* a key is read, so an unusable base URL never causes a
credential to be looked up. The message renders the URL without its userinfo
or query, so it is safe to log. `docs/user/models-and-providers.md` gains a
**Base URLs** section stating all of it. (REV-2 A.6.)
- `baikai-openai`, `baikai-claude`: the `ClientEnv` cache was duplicated in each
package and keyed on the raw base-URL text, so `https://h` and `https://h/`
were two TLS managers and two connection pools to one host. There is now one
cache, in `Baikai.Http`, keyed on the canonical rendering of the parsed base
URL. `Transport.getClientEnvCached` and `Transport.cachedClientEnvCount` are
re-exports of the core functions and keep their signatures.
- `baikai-openai`: the Codex interactive launcher now **refuses the two approval
policies the installed CLI rejects**. `codex --help` at `codex-cli 0.149.1`
lists exactly `on-request` and `never` for `--ask-for-approval`;
`CodexApprovalUntrusted` and `CodexApprovalOnFailure` are older spellings the
CLI answers with `error: invalid value 'untrusted' for
'--ask-for-approval'`. Rendering them made a launch return `Right` carrying a
non-zero exit code — a session that ran and failed — instead of the `Left
SafetyNotExpressible` this module promises for a policy that cannot be
honoured. They are refused before any process is created, and refused rather
than quietly mapped onto `on-request`, because substituting a different
approval policy would change what the caller asked for. The constructors and
their spellings are unchanged, so code that matches on `CodexApprovalPolicy`
keeps compiling.
## [baikai-trace-otel 0.4.0.0] - 2026-08-28
### Added
- `baikai-trace-otel`: `OtelSinkOptions` derives `Generic`, so `#spanName`
resolves on it. No `Eq` or `Show`: `OpenTelemetry.Context.Context` has neither,
and an instance that ignored `parentContext` would be a lie. (REV-2 G.6.)
- `baikai-trace-otel`: `OtelSinkOptions.parentContext :: Maybe Context`, default
`Nothing`. When set, every span the sink opens becomes a child of the span in
that context instead of a root, so a call can be nested under the caller's own
request span. It is a value fixed when the sink is built rather than an action
run per call, because the fold runs on baikai's trace worker thread where the
caller's thread-local context is invisible: capture the context on your own
thread (`ctx <- getContext`, or `Context.insertSpan mySpan Context.empty`) and
build the sink for that request. __Breaking for positional construction__ of
`OtelSinkOptions`; the documented path is a record update on
`defaultOtelSinkOptions`. (REV-2 D.9.)
### Changed
- `baikai-trace-otel`: the `baikai.evidence.strength` span attribute is rendered
by `Baikai.Evidence.renderEvidenceStrength`, the function the JSON encoding
uses, instead of a second spelling local to the sink that could drift from it.
- `baikai-trace-otel`: `gen_ai.response.model` is set only by the evidence
branch, from the model the provider reported. The terminal branch set it from
the *requested* id, and since evidence is pushed before the terminal and
`addAttributes` replaces a key, that both labelled a request as an observation
on every call without evidence and overwrote the genuinely observed value on
every call with one. (REV-2 D.1.)
## [baikai-effectful 0.4.0.0] - 2026-08-28
### Changed
- `baikai-effectful` (breaking): the version is a **major** bump although this
package's own exports are unchanged. Its `baikai` bound moves to `^>=0.6.0`,
and the `Baikai` effect's three operations are typed in `Model`, `Context`,
`Options` and `Response` — every one of which baikai 0.6.0.0 changes
breakingly. A consumer therefore meets a break through this package even
though nothing in it was renamed, so the number says so rather than making
`0.3.0.4` look like a safe upgrade.
- `baikai-effectful`: no longer depends on `streamly`. Both stanzas listed it
while every module imports only `Streamly.Data.Fold` and
`Streamly.Data.Stream`, which are `streamly-core`. (REV-2 minor.)
## [baikai-kit 0.2.0.0] - 2026-08-28
### Added
- `baikai-kit`: `Baikai.Kit.Error` with the closed `KitError` sum, its
`Exception` instance and `renderKitError`; `Baikai.Kit.Path.safeSourcePath`,
which resolves an untrusted relative source below the kit checkout and refuses
a symbolic link in any component or a canonical path outside the checkout;
`Baikai.Kit.Manifest.itemSources`/`ItemSources`, the one pure derivation of an
item's source list, and `supportedManifestVersions`;
`Baikai.Kit.Sidecar.hashEntries`; `Baikai.Kit.Repo.KitRepo`/`RepoRefresh`;
`Baikai.Kit.Install.installFrom`, `renderAvailable` and `UpdateReport`;
`Baikai.Kit.Status.StatusReport`, `UpstreamAvailability` and the now-pure
`renderStatusTable`; `Baikai.Kit.Command.runKitCommand`. `KitState` gains
`KitUpstreamRefused`, rendered `refused`. (REV-2 E.5, F.10, F.11.)
- `baikai-kit`: `Baikai.Kit.Install.OverwritePolicy` (`KeepLocalEdits`,
`OverwriteLocalEdits`), `reinstallPresent` (the network-free half of
`updateKit`), and `PlannedWrite`/`WriteContent`/`executePlan`/`executePlanWith`
as a test seam. `SidecarMeta` gains `installedFiles` and `installedHash`,
which record what this tool wrote for one provider and the hash of exactly
those bytes; `newSidecarMeta` takes both. `kit update` gains `--force`.
(REV-2 F.12, Theme 8.2.)
### Changed
- **Breaking.** `baikai-kit`: every library function returns
`Either KitError a` and prints nothing; only
`Baikai.Kit.Command.runKit` prints `Error: …` and exits 1. `loadManifest`,
`loadManifestMaybe`, `installItem`, `listAvailable`, `uninstallItem`,
`updateKit` and `ensureKitRepo` change shape accordingly, `computeKitHash`
takes the kit root, a base and relative file names, `kitStatus` returns a
`StatusReport` instead of printing, and `KitUpdate`'s report is rendered by
the caller. See `docs/adr/0013-library-code-never-calls-exitfailure.md`. A
consumer that only calls `runKit` and `kitCommandParser` needs no change; one
that calls the library directly binds `Right`. (REV-2 F.11.)
- `baikai-kit`: a kit is plain files. Install, the content hash and `kit status`
resolve every listed source through `safeSourcePath`, so a kit repository that
commits a symbolic link can no longer have a file read through it and copied
into a provider directory. `kit status` shows such an item as `refused`.
(REV-2 E.5 = F.10.)
- `baikai-kit`: a manifest whose `version` is not 1 or 2 is refused with
`KitManifestVersionUnsupported` instead of being decoded and installed.
(REV-2 F.12.)
- `baikai-kit`: an agent that lists several `files` installs all of them. The
first becomes the provider's agent file as before, and each remaining file
goes into a resource directory named after the agent beside it
(`<agents dir>/<name>/<file>`), which uninstall removes with the agent. Only
the first file used to be installed. (REV-2 F.12.)
- `baikai-kit`: `kit update` skips an item whose installed files no longer hash
to what its sidecar recorded, printing the `--force` invocation that would
overwrite them; `kit update --force` reinstalls anyway. Sidecars written
before this release carry no such hash and are updated without the check.
(REV-2 Theme 8.2.)
### Removed
- **Breaking.** `baikai-kit`: `Baikai.Kit.Path.safeUnder` (exported and unused),
`Baikai.Kit.Manifest.agentSources` (replaced by `itemSources`) and
`Baikai.Kit.Install.uninstallOutcomes` (absorbed by `uninstallItem`, which now
returns the outcomes for the caller to render). The internal `requireSafe` and
`Baikai.Kit.Status.resolveCacheOrEmpty` are gone with the exits they wrapped.
### Fixed
- `baikai-kit`: `kit status` with no cache and no network prints
`No kit items installed.` and exits 0. It used to exit 1: the guard around
`ensureKitRepo` caught `IOException`, which is not what `exitFailure` throws.
(REV-2 F.11.)
- `baikai-kit`: `Baikai.Kit.Status.upstreamHash` joined the manifest `path`
without validating it, a second unsanitised join that grew after the July
hardening pass validated the first. Both now go through `itemSources` and
`safeSourcePath`. (REV-2 Theme 8.1.)
- `baikai-kit`: an install that fails while renaming files into place now
restores what was there before, or names the paths it could not restore.
Phase two was a bare loop of renames, so a failure part-way left earlier
renames in place while the message said "no changes were made". Temporary
files are also created with `openTempFile`, so two concurrent installs of one
item no longer clobber each other's staging file, and a destination that is a
directory is refused before anything is written. (REV-2 F.12.)
- `baikai-kit`: `Baikai.Kit.Install.stripYamlFrontmatter` normalises line
endings to LF on every branch. Input without frontmatter, and input whose
frontmatter is never closed, used to keep their `\r` characters and leak them
into the Codex agent TOML. (REV-2 Theme 8.7.)
- `baikai-kit`: an `IOException` raised while reinstalling during `kit update`
is returned as `KitWriteFailed` instead of escaping as an uncaught exception.
(REV-2 Theme 8.4.)
## [baikai-agent 0.2.0.0] - 2026-08-28
### Added
- `baikai-agent`: three operator-only `policy` keys — `policy.allowed-tools`,
`policy.max-timeout` (a duration or `"unlimited"`) and
`policy.max-output-limit` (a byte count or `"unlimited"`) — each defaulting
from `defaultAgentCeiling`, and all six ceiling fields now printed by
`agent show` and carried in its `--json` object.
- `baikai-agent`: `Baikai.Agent.Config.repositoryScopeViolations`, which reads
the resolution report to say which values the untrusted repository file was
not allowed to supply at all. `Baikai.Agent.Cli` concatenates its answer with
the pure ceiling's, so an operator sees one refusal naming every problem.
### Changed
- `baikai-agent` (breaking): `AgentConfigScope`'s constructors are
`AgentUserScope` and `AgentRepositoryScope`. `UserScope` collided with
`baikai-kit`'s `KitScope` constructor of the same name, the one clash between
two baikai-family packages. (REV-2 G.5.)
- `baikai-agent` (breaking): a relative `working-dir` resolves against the
repository root rather than the process's own directory, so `working-dir "."`
means the checkout whichever file declared it. Resolving against the process
directory made `"."` mean two places when two documents defined one job, since
which one it was depended on which layer won. An absolute path is unchanged.
(REV-2 F.14.)
- `baikai-agent` (breaking): every `--json` output is now built with `aeson`
rather than a hand-rolled writer, and `agent show --json` always emits one
object with the same seven keys — `job`, `outcome` (`shown`, `refused` or
`failed`), `exitCode`, `message`, `configuration`, `ceiling`, `command` —
with `null` for the parts that do not apply. Previously a refusal emitted a
different shape from a success and a document that would not parse emitted a
bare resolution report or nothing at all, so a reader had to know which
failure mode it was looking at before it could find the exit code. `run --json`
keeps its `outcome` values and `list --json` is unchanged. (REV-2 F.14.)
- `baikai-agent` (breaking): `--run-id` or `--require-evidence` without either
`--evidence-file` or `--json` is now a usage error (64) naming both fixes.
Before, the record was built — a `--version` probe of the tool and two digests
— and then dropped. Under `--json` the record now travels in the envelope as
`evidence`, encoded by the same `ToJSON` `--evidence-file` writes.
- `baikai-agent`: `agent show` and `agent run` no longer print another job's
unknown-key warnings, or the operator file's `policy` keys. The declaration
describes one job and the ceiling is a separate declaration, so `settei` warns
about both; neither is a mistake and a document with four jobs printed three
jobs' worth of noise on every run. A misspelled key inside the selected job
still warns, and a `policy` node in the *repository* document earns exactly one
notice saying it has no effect. `Baikai.Agent.Config` exports the two filters,
`relevantWarnings` and `repositoryPolicyNotice`. (REV-2 F.13.)
- `baikai-agent`: an evidence record's `endpoint` resolves a relative executable
against the job's working directory before probing it, because that is what
the child execs. A job whose `executable` is `./bin/agent` previously reported
a path resolved against the parent's own directory, which does not exist.
`Baikai.Agent.Run` exports `executableForEvidence`. (REV-2 F.13.)
- `baikai-agent`: a failed run's `error_info.message` keeps the last
`errorInfoStderrTailBytes` (4096) bytes of standard error, prefixed with how
many earlier bytes were dropped, instead of the whole captured stream — which
the output limit allows to reach four mebibytes by default. `Baikai.Agent.Run`
exports the constant. (REV-2 F.13.)
- `baikai-agent`: `--evidence-file` stages through a uniquely named temporary
file created with `O_EXCL` beside the destination, instead of the destination
plus `.partial`. A symbolic link planted at the old, guessable name was
followed, which let an unattended run overwrite a file of the planter's
choosing. (REV-2 F.13.)
- `baikai-agent` (breaking): an operator configuration file that lies inside the
repository root is refused with exit 78, naming the file and the root, and no
ceiling is established. The source list already refused the repository
*document*; this closes the shape where the repository supplies the *operator*
document, which both `--user-config .baikai/policy.kdl` and
`XDG_CONFIG_HOME=$PWD/.baikai` produce. `--user-config`, `XDG_CONFIG_HOME` and
`HOME` remain the operator's own inputs: the ceiling is exactly as trustworthy
as the process environment that selects it, and the guide now says so.
(REV-2 F.4.)
- `baikai-agent` (breaking): an unrecognised key under the operator file's
`policy` node is an error rather than a warning, naming the file and every
such key. Everywhere else a forward-compatible file should not stop an older
binary; under `policy` a misspelling would silently leave the default ceiling
in force, which for the one node whose purpose is limiting authority is
indefensible. Two `AgentConfigError` constructors are added,
`CeilingFileInsideRepository` and `UnknownPolicySetting`.
- `baikai-agent` (breaking): `AgentConfigPaths` gains `repositoryRoot`, the
directory the process runs in. `--config PATH` chooses which file supplies
repository-scope settings and does not move the root, because the root is what
confines a repository-supplied `working-dir`.
- `baikai-agent` (breaking): a repository configuration file may no longer set
`executable` or a non-empty `extra-dirs`, and its `working-dir` must resolve —
after following symbolic links — inside the repository root. Each is refused
with exit 77 naming the setting, or naming both directories. The operator's
own file and `--set` may still set all three. `executable` turns configuration
into code execution with the operator's environment and the prompt on standard
input; `extra-dirs` inside the root adds nothing the working directory does not
already give, so the only ones a checkout would ask for are outside it.
(REV-2 F.3.)
### Removed
- `baikai-agent` (breaking): the `BAIKAI_AGENT_EXECUTABLE` environment binding.
An environment variable is inherited by every child process and is easy to set
by accident, and naming the program to run is the widest widening there is.
An operator whose installation is not on `PATH` writes `executable` in their
own configuration file or passes `--set`.
### Fixed
- `baikai-agent`: a timed-out run now **escalates to `SIGKILL`**. The runner
interrupts the child's whole process group, then terminates it, then kills it,
each of the first two stages bounded by the grace period and ended early once
the leader has been reaped and no member of the group is left. Previously the
last resort was `terminateProcess` followed by an *unbounded* wait, so a
coding agent that ignored `SIGTERM` — or a grandchild holding the output pipe
— hung the run for as long as it chose to live, with the deadline already
past. Polling the group rather than waiting on the leader alone is also what
gives a grandchild the same grace the agent gets.
- `baikai-agent`: a timed-out run **reports the output it drained**. `baikai
agent run` prints it under the same stream discipline a finished run gets, so
`response=$(baikai agent run job)` under `capture` receives the partial answer
with `$?` set to 75, and `--json`'s failure envelope carries the same
`stdout`, `stdoutTruncated`, `stderr` and `stderrTruncated` fields. A drain
interrupted because something outside the process group still held the pipe
open keeps its bytes too, reported as truncated.
- `baikai-agent`: the `baikai` command writes its output as **UTF-8 bytes**
rather than through the locale encoding. Where an unattended run actually
happens — cron, a systemd unit, a container — the environment says `LANG=C`,
and on a platform whose locale encoding follows it a single accented character
in the agent's answer made the write throw after the run had already finished:
exit 1, answer lost. This mirrors what the prompt read and the prompt write
have always done.
- `baikai-agent`: the `baikai` executable now links the **threaded runtime**
(`ghc-options: -threaded` on the `executable baikai` stanza). Without it a
blocking operating-system call — the `waitpid` inside
`System.Process.waitForProcess` — stopped every Haskell thread in the
installed binary, so a job's configured `timeout` could never fire and a
coding agent that wrote more than one pipe buffer deadlocked against the
runner's drain threads. Both defects existed only in the shipped executable:
the test suite was already compiled `-threaded`, so every runner test passed
under a runtime the binary did not have.
The suite now proves the runtime the binary ships with rather than its own.
`baikai-agent/test/BinaryTests.hs` spawns the built executable — cabal builds
it first and puts it on the suite's `PATH` through
`build-tool-depends: baikai-agent:baikai` — asserts that `baikai +RTS --info`
reports `rts_thr`, and runs `baikai agent run` against a stub agent that
outlives its deadline, requiring exit 75 within seconds and the whole process
group gone. See
[docs/adr/0006](docs/adr/0006-a-process-spawning-executable-ships-on-the-threaded-runtime.md).
## [baikai 0.5.0.0] - 2026-08-05
### Added
- `baikai`: new exposed module `Baikai.Agent`, the provider-neutral vocabulary
for an **unattended coding-agent run** — a run with no terminal and no human,
which owns its own tool loop, may change files inside directories the caller
authorized, and returns a process result rather than a `Response`. It defines
`AgentRunRequest` (with a required `workingDir`), `AgentRunResult`, the
`AgentCapability` profile (`read-only`, `edit-workspace`, `full-access`),
`AgentSafety`, the `AgentOutputMode` and `AgentCapturedOutput` output
discipline, the `AgentCommand` renderer/runner boundary with an explicit
prompt transport, and the `AgentRenderError` / `AgentRunFailure` taxonomies.
- `baikai`: the operator policy ceiling — `AgentCeiling`,
`defaultAgentCeiling`, `CeilingViolation`, and the pure `applyAgentCeiling`.
It returns a request unchanged when it is within the ceiling and reports
every violation when it is not; it never clamps an over-broad request to the
permitted value. The default ceiling permits read-only and edit-workspace
authority and refuses full access and raw provider arguments.
`Baikai.Agent` itself is vocabulary and pure policy algebra only: it spawns no
process and renders no command-line flags. Those live in the vendor packages
and in `baikai-agent`, below. The module is deliberately not re-exported from
the umbrella `Baikai` module, because its field accessors share names with
`Baikai.Interactive`, so `import Baikai` continues to compile unchanged.
- `baikai`: new exposed module `Baikai.Evidence`, the vocabulary for
**verifiable model-call evidence** — a record of what actually crossed the
boundary to a provider, as opposed to what the process was configured to ask
for. It defines `ModelCallEvidence` and the `evidenceSchemaVersion` string
consumers pin against, `Observed` (a deliberate non-`Maybe` for a value the
provider either did or did not report, with no function that supplies a
default), `ThinkingTranslation` with its `ThinkingMode` and
`ThinkingAdjustment` enumerations describing what a requested
reasoning-effort level actually became on the wire and every clamp, collapse,
or drop applied on the way, `EndpointIdentity` and `TransportKind`,
`CallStatus`, and the ascending `EvidenceStrength` scale.
It also provides the canonical hashing core: `canonicalEncode` gives a JSON
value exactly one byte representation (object keys sorted, no insignificant
whitespace, numbers normalised so `1`, `1.0`, `1.00`, and `1e0` all encode as
`1`, and a hand-written string escaper so an aeson upgrade cannot silently
invalidate a recorded digest); `commitmentDigest` hashes a full request
envelope, and `configurationDigest` hashes an allow-list projection
(`configurationProjection`) that keeps configuration and replaces content with
structural summaries, so two calls that ask the same model the same way about
different subjects agree. The two digests are separate on purpose: the first
binds a record to a particular request, the second is safe to compare across
runs that legitimately differ in content.
Nothing constructs a `ModelCallEvidence` from a real call yet, and no existing
behaviour changed. New dependencies: `cryptohash-sha256` and
`base16-bytestring`, both single-purpose packages chosen over a full
cryptographic framework.
- `baikai`: `Options` gains an `evidence` field carrying an optional
`EvidenceRequest` — the caller's run identifier, retry provenance, and how
strictly they need evidence. A call whose `evidence` is `Nothing`, which is
every call that does not opt in, behaves exactly as it did before: no digest
is computed and no evidence is emitted.
- (Entry added 2026-08-27; the behaviour shipped in 0.5.0.0.) `baikai`: **strict
evidence mode**. `EvidenceStrictness` is `EvidenceBestEffort` or
`EvidenceRequired !EvidenceStrength`, and a caller who asks for the second
gets a call that **refuses to start** — before any request is built or any
connection opened — when the configuration cannot reach the strength asked
for: `Baikai.Evidence.Build.checkEvidenceRequirements` compares the
requirement against what the provider can deliver and against the thinking
translation, and `completeRequest` / `streamRequest` return an error-shaped
response or a terminal `EventError` instead of dispatching. The gate is
pre-dispatch by design; that is the only point at which refusing is still
free.
- (Entry added 2026-08-27; the behaviour shipped in 0.5.0.0.) `baikai`:
**sink-failure semantics under strict mode**. `Baikai.Evidence.Build`
exports `onSinkFailure`, `sinkFailureIsFatal` and `sinkFailureError`: a trace
sink that throws fails an `EvidenceRequired` caller's call, because a record
the sink did not confirm written is not a record, while a best-effort caller's
call succeeds with the failure reported on stderr.
- (Entry added 2026-08-27; the behaviour shipped in 0.5.0.0.) **Breaking.**
`baikai`: `Baikai.Provider.Registry.ApiProvider` gained a fourth field,
`describeThinking :: Model -> Options -> ThinkingTranslation`, which the
pre-dispatch strictness gate calls to learn what a provider would do with the
caller's reasoning-effort request without sending anything. Every third-party
provider constructed with the `ApiProvider` constructor stopped compiling.
This was not recorded at the time; it is the defect that made 0.6.0.0 hide the
constructor behind `apiProvider` so that the next field addition is a minor
release.
- `baikai`: model-call evidence is now **produced and emitted**. A caller who
sets `Options.evidence` gets exactly one `call_evidence` line per call from
their trace sink, under every way a call can end: success, provider failure, a
consumer that abandons the stream (status `aborted`, not `failed` — an abort
is the consumer's doing and reporting it as a provider failure would
misattribute it), and dispatch that found no registered handler.
New exposed module `Baikai.Evidence.Build` bridges the vocabulary to the
`Model` and `Options` records: `minimalEvidence` and `prepareEvidence` build a
record, `dispatchEnvelope` supplies the request envelope for the paths where
no adapter ran, `sanitizeEndpoint` reduces a base URL to scheme/host/port/path
with the query string and any userinfo dropped wholesale, and `onSinkFailure`
is the hook a future release replaces to make a strict caller's call fail when
the trace sink does.
Every record this release produces has `strength` `requested_only` and every
provider-observed field set to `"unobserved"`. That is not a placeholder: it
is a truthful record for a transport that has not yet been taught to observe
anything. Later releases teach each transport to observe more.
(Correction added 2026-08-27: the two paragraphs above describe the release
inaccurately and are kept as shipped rather than rewritten. `onSinkFailure`
did not await a future release — it shipped in 0.5.0.0 together with
`sinkFailureIsFatal` and `sinkFailureError`, which already fail a strict
caller's call when the sink throws. And not every 0.5.0.0 record has `strength`
`requested_only`: the provider entries below describe what each transport
reports, and the HTTP adapters reach `correlated` and `model_observed`.)
**A caller who does not opt in pays nothing.** With `Options.evidence` absent
no digest is computed, no call identifier is generated, no evidence event is
emitted, and the request envelope is never even forced — the gate lives inside
the shared builder rather than at each adapter's call site, and the envelope
parameter is deliberately lazy. Both facts are guarded by tests.
- `baikai`: `TraceEvent` gains a `CallEvidence` constructor, encoded as
`{"kind":"call_evidence", …}`. A consumer whose pattern match over `TraceEvent`
is exhaustive must add a branch; one with a wildcard is unaffected. Filter for
it with `jq 'select(.kind == "call_evidence") | .evidence'`. Note that a trace
line carries its fields alongside the `kind` discriminator rather than nested
under a `data` key, and that the evidence record inside spells its own fields
in snake_case — the two encodings differ deliberately, because an evidence
record must render an absent field as explicit `null` while a trace line drops
it to stay small.
- `baikai`: `Baikai.Provider.Cli.Internal` — the module the two subprocess
providers share — gains the vocabulary for reading what a coding-agent CLI
reported about its own run. `CodexRunReport` and the new
`parseCodexJsonlStream :: Stream IO ByteString -> IO CodexRunReport` fold the
`codex exec --json` event stream into its assistant text, its thread
identifier, and its token counts, instead of concatenating agent-message text
and discarding everything else. `ClaudeCliReport` and
`decodeClaudeCliResult` do the same for `claude -p --output-format json`.
Every field but the message text is optional, because both tools' event
schemas have changed across versions and an absent field is a genuine absence
rather than a parse failure. **Breaking** for anyone calling
`parseCodexJsonlStream` directly: its result type is no longer `Text`. This is
an internal module and is documented as outside the PVP guarantee.
- `baikai`: `Baikai.Provider.Cli.Internal` also gains `ExecutableIdentity` and
`executableIdentity`, which resolve a configured executable name to an
absolute path and read the tool's own `--version` line. The probe is cached
per resolved name for the lifetime of the process, because spawning it per
model call would roughly double the process cost of the cheapest possible
call, and it is bounded by a five-second timeout so a tool that hangs on
`--version` cannot wedge a model call. (Corrected 2026-08-27: the entry said
two seconds; `versionProbeMicros` has always been five.) A probe that fails records the version
as absent rather than failing the call. It is only ever called from inside
the evidence branch: a caller who asked for no evidence must not pay for a
process whose only purpose is to describe a tool they were about to run
anyway.
- `baikai`: `subprocessStrength` and `cliResponseEnvelope`, also in
`Baikai.Provider.Cli.Internal`. The former derives a subprocess call's
evidence strength from what the tool reported and **nothing else** — the exit
status is deliberately not one of its arguments. The latter spells the
response-commitment envelope with the same three keys, in the same shapes, as
the two API transports build by hand, so a verifier holding a response can
recompute the digest without first knowing which transport served it.
- `baikai`: `Baikai.Agent` gains `AgentRunOutcome` and `agentRunOutcome`. It
pairs what an unattended run did — the existing
`Either AgentRunFailure AgentRunResult` — with the evidence the runner built
for it. The evidence is a sibling of the outcome rather than a field on
`AgentRunResult` because the run that most needs a record is one that did not
produce a result: a run killed by its own timeout reports
`Left (RunTimedOut …)`, so a record hanging off the `Right` would be
unreachable exactly there.
### Fixed
- `baikai`: a `call_evidence` event is now emitted **before** its call's
terminal `call_finished` or `call_failed`, rather than after. The
OpenTelemetry sink ends and removes a call's span on the terminal, so under
the old order its evidence-attribute branch was unreachable from any real
call and every backend saw a span with no evidence on it — nothing failed,
the attributes were simply never there. No consumer can have depended on the
old order, because no consumer has ever seen a `call_evidence` line.
- `baikai`: the `ThinkingFormatOpenAI` Haddock in `Baikai.Compat` listed the
native `reasoning_effort` vocabulary as `minimal | low | medium | high`, which
predates `xhigh` and `max`. It now lists all six and states that this shape
alone sends the canonical baikai level verbatim while the other six clamp
through `compatibleEffort`. No behaviour changed: the native path's exclusion
from that clamp is deliberate and is guarded by two named tests in
`baikai-openai/test/ShapeSpec.hs`. A reader who consulted the comment to
decide whether `xhigh` was safe to use against OpenAI has until now been told
something untrue.
### Changed
- **Breaking:** `baikai`: `TerminalPayload` gains an `evidence` field and the two
terminal smart constructors take it as their new first argument:
`doneTerminal :: Maybe ModelCallEvidence -> Maybe Text -> StopReason -> Message -> TerminalPayload`
and `errorTerminal` likewise. `Response` gains the same field. A custom
provider implementation must pass `Nothing` (or a record it builds through
`Baikai.Evidence.Build`); a custom `Response` built with the record
constructor must add `evidence = Nothing`. Code that only pattern-matches on
these types is unaffected.
- **Breaking:** `baikai`: `CallFinished` gains `cachedInputTokens`,
`cacheWriteTokens`, `reasoningTokens`, and `totalTokens`. The trace path used
to drop counts that `Baikai.Cost.Log.CallLogEntry` kept from the same `Usage`
value, which made the cost log strictly more faithful than the trace.
- **Breaking:** `baikai`: a computed cost of **zero is now reported as zero**
rather than suppressed, in `CallFinished` and at all three `CallLogEntry`
construction sites. Previously `usd` was omitted whenever the cost came out at
zero, so "this call was free" and "baikai could not price this call" were
indistinguishable — and the subscription-based CLI providers always price at
zero, so that was the common case rather than a corner. **A cost dashboard
that treated an absent `usd` as "unpriced" will now count those calls as
costing zero.** That is the correct reading, but it changes what such a
dashboard shows.
- **Breaking:** `baikai`: `FromJSON TraceEvent` is written out by hand instead of
derived. The three pre-existing kinds decode exactly as before; a
`call_evidence` line fails to parse with a message saying to read it as a
plain `Data.Aeson.Value`. `ModelCallEvidence` has no `FromJSON` on purpose —
it embeds a `Cost` whose exact `Rational` amounts encode through an
approximating `Scientific`, so a decoder would return a different value than
was encoded — and manufacturing that fidelity would be the precise failure
this vocabulary exists to eliminate.
- `baikai`: `Baikai.Trace.Sink.renderHuman` renders a `CallEvidence` event as a
single `EVIDENCE run=… call=… strength=…` line rather than the whole record. A
human-readable sink is for watching calls go by; the full record is meant to
be read out of `fileSink` output by a machine.
- `baikai`: call identifiers on the trace path are now globally unique.
`Baikai.Evidence.newCallId` produces 32 lowercase hexadecimal characters
carrying 128 bits — 48 bits of Unix time in milliseconds, 48 bits of a
per-process random seed drawn once from `/dev/urandom`, and a 32-bit counter.
The previous generator combined the process-start *second* with a
process-local counter into 16 characters, so two processes started within the
same second emitted identical identifier sequences; its own documentation
claimed only per-process uniqueness. Identifiers still sort chronologically
and are still not secrets.
`Baikai.Trace.newEventId` keeps its name and signature, delegates to
`newCallId`, and is now deprecated. Anything that pinned the 16-character
width — a log parser, a fixture, a column type — must widen to 32.
- `baikai`: `renderCeilingViolation` no longer prints the raw provider arguments
a `ProviderArgsForbidden` violation carries. It reports how many were
requested and states that their values are not shown. Raw provider arguments
are the one part of a job description that can hold a credential — the
configuration layer classifies the setting secret for that reason — and a
refusal message that quoted them defeated the classification. The constructor
keeps its `[Text]` payload so a programmatic caller can still inspect it.
## [baikai-claude 0.5.0.0] - 2026-08-05
### Added
- `baikai-claude`: new exposed module `Baikai.Provider.Claude.Agent` with
`ClaudeAgentConfig`, `defaultClaudeAgentConfig`, and `claudeAgentCommand`, a
pure renderer from an unattended `AgentRunRequest` to the `claude` argument
vector. It maps the capability profile onto `--permission-mode`
(`plan` / `acceptEdits` / `bypassPermissions`), joins a tool allow-list into
one `--allowedTools` argument, repeats `--add-dir` per extra directory, always
emits `-p`, and emits `--no-session-persistence` unless `persistSession` is
set. The prompt travels on standard input and appears nowhere in the argument
vector. A request naming a different provider is refused with
`ProviderMismatch`. Nothing is spawned.
- `baikai-claude`: the Anthropic Messages provider now fills in the evidence
record it previously left blank. It records the model **Anthropic reported
running** (read from the `message_start` event, which the adapter already
decoded for the response id and then discarded), Anthropic's `request-id`
correlation header, the response id, the token counts Anthropic actually
reported, and a commitment digest over the assembled response. A field the
provider did not report stays `"unobserved"` and is never backfilled from the
request — in particular, a stream that fails before `message_start` reports no
observed model at all. `strength` is `model_observed` when both the model and a
correlation identifier arrived, `correlated` when only the identifier did, and
`requested_only` otherwise; a 2xx status never raises it, because a 200 means
the request was accepted, not that any particular model ran.
`fully_observed` is unreachable on this transport, since Anthropic does not
echo the thinking configuration it applied.
- `baikai-claude`: an evidence record's `thinking` field now describes what the
caller's reasoning-effort preference actually became on the wire, including
three downgrades that were previously invisible everywhere in baikai's output:
asking for thinking on a model that does not advertise `reasoning`
(`thinking_dropped_unsupported_model`); asking for a level whose token budget
does not fit under the resolved output-token ceiling
(`thinking_dropped_budget_exceeded`, carrying both colliding numbers), which is
reachable by lowering `maxTokens` alone; and asking for `high` on an
adaptive-thinking model, which sends no effort field and so is
wire-indistinguishable from taking Anthropic's default depth
(`effort_omitted`). `minimal` on an adaptive model reports `effort_clamped`,
because Anthropic's adaptive vocabulary has no `minimal`.
- `baikai-claude`: new exports from `Baikai.Provider.Claude.Sse` —
`ResponseMetadata` and `capturedHeaderNames` — and from
`Baikai.Provider.Claude.Api` — `claudeMessagesStreamWith`, `SseDriver`, and
`anthropicStrength`. Response-header capture is an **allow-list**
(`request-id`, `x-request-id`, `cf-ray`, in that preference order), not a
denylist, so a header a future gateway adds is not recorded by default.
- `baikai-claude` and `baikai-openai`: both subprocess providers now fill in the
evidence record they previously left blank, and both export the translation
function that describes it — `claudeCliThinking` and `codexCliThinking`. They
record the session or thread identifier the tool reported, the token counts it
reported, the model it named when it names one, the resolved executable path
in place of an endpoint URL, the tool's own `--version` string as the
implementation version (for this transport the tool *is* the implementation),
a request commitment over the rendered argument vector, and a response
commitment over the assembled answer.
**A zero exit status never raises the strength.** A coding-agent CLI that
exits zero has demonstrated that it ran and did not crash; it has not stated
which model served the request. Subprocess calls almost always exit zero, so
encoding that as corroboration would make the weakest evidence in the system
look like the strongest. `strength` is `model_observed` only when the tool
named both an identifier and a model, `correlated` when it named only an
identifier, and `requested_only` otherwise.
The two transports differ in how far they can get. `claude` names the model
that consumed tokens in its result event's `modelUsage` map, complete with a
context-window variant marker such as `[1m]`, so a Claude CLI run can reach
`model_observed`. `codex-cli 0.146.0` names no model anywhere in its event
stream, so **no** Codex CLI run can exceed `correlated` — backfilling the
`--model` flag baikai passed would report the request as an observation.
- `baikai-claude`: an evidence record's `thinking` field now describes what a
reasoning-effort request became on the `claude` command line: mode `flag`,
wire field `--effort`, and an `effort_clamped` adjustment recording the
`minimal` → `low` collapse, because the tool's `--effort` flag has no
`minimal`. A caller asking for `minimal` and a caller asking for `low` produce
byte-identical argument vectors — and therefore identical request commitment
digests — so the translation is the only place that difference survives.
- **Breaking:** `baikai-claude` and `baikai-openai`: `claudeAgentCommand` and
`codexAgentCommand` return `(AgentCommand, ThinkingTranslation)` rather than
`AgentCommand`. The runner deliberately imports no vendor renderer, so it
cannot derive the translation and has to be handed it. A caller that only
wants the command writes `fmap fst`. Both modules also export the translation
function alone — `claudeAgentThinking` and `codexAgentThinking` — for asking
what a level would become without rendering anything.
### Fixed
- **Loud:** `baikai-claude` and `baikai-openai`: both subprocess providers
hardcoded `usage = zeroUsage` on every call, so a cost dashboard saw every
`claude -p` and `codex exec` call as consuming no tokens and costing nothing.
Both tools report their own token counts and baikai now carries them through,
normalized into the disjoint `Usage` convention: `claude`'s counts are
Anthropic-shaped and already disjoint, while `codex` reports OpenAI-style
inclusive prompt counts, so its cached tokens are subtracted out of
`inputTokens`. `claude` additionally reports a `total_cost_usd`, which now
populates `Usage.cost` exactly rather than being reported as zero.
**A dashboard that read these calls as free will now see real tokens and, for
`claude`, a real cost.** That is the correction, not a regression — but it
changes what existing reports show, and totals over historical data will not
match totals over new data.
- `baikai-claude`: `Response.responseId` was always `Nothing` on the `claude -p`
transport even though `ClaudeCliResult` decoded the tool's `session_id` one
screen earlier and then dropped it. It now carries that identifier, on both
the successful and the failed terminal. `baikai-openai`: the same for
`codex exec`, whose thread identifier was filtered out of the event stream
along with everything that was not an `agent_message`. These are the handles
each vendor's support tooling looks a run up by.
### Changed
- **Breaking:** `baikai-claude`: `Baikai.Provider.Claude.Sse`'s four streaming
entry points — `claudeSseStream`, `claudeSseStreamValue`,
`claudeSseStreamValueWithHeaders`, and `sseFromResponse` — take a new
`ResponseMetadata -> IO ()` callback immediately before the existing per-event
callback. It fires exactly once, before the first event, on both the success
and the non-2xx path. Pass `(\_ -> pure ())` to keep the previous behaviour.
The callback is separate rather than a widening of the per-event one because
the per-event callback runs once per SSE frame and response-level data does not
belong on that path.
- **Breaking:** `baikai-claude`: `Baikai.Provider.Claude.Internal.Request`'s
`mapRequest` now returns
`Either Text (Messages.CreateMessage, ThinkingTranslation)` and
`computeThinking` returns `(ThinkingPlan, ThinkingTranslation)`. Take `fst` to
keep the previous value. This module is exposed for provider tests and
debugging and its header states it is not covered by PVP compatibility
guarantees, but the change is recorded here because that is not a licence to
break a consumer silently.
- **Breaking:** `baikai-claude`: `claudeInteractiveCommand` now returns
`Either AgentRenderError (FilePath, [String])` and `launchClaudeInteractive`
returns `IO (Either AgentRenderError InteractiveLaunchResult)`. A request
whose `safety` is a `CodexSandbox` policy — which Claude Code cannot express
— is refused with `SafetyNotExpressible AgentClaude`, naming the rejected
sandbox mode and approval policy and suggesting `ClaudeAllowedTools` or
`DefaultSafety`. Previously the policy was silently discarded and an
**unrestricted** Claude session was started and reported as a success. A
`Left` means no process was started; a `Right` with a non-zero exit code
means the session ran and exited non-zero. `DefaultSafety` and an empty
`ClaudeAllowedTools` list still render no safety flag and are never refused,
and no previously rendered argument vector changed. Callers must handle the
refusal branch.
## [baikai-openai 0.5.0.0] - 2026-08-05
### Added
- `baikai-openai`: new exposed module `Baikai.Provider.OpenAI.Agent` with
`CodexAgentConfig`, `defaultCodexAgentConfig`, and `codexAgentCommand`, the
same renderer for `codex exec`. It maps the capability profile onto
`--sandbox` (`read-only` / `workspace-write` / `danger-full-access`), emits
`--cd` for the working root, and defaults `--skip-git-repo-check` and
`--ephemeral` on. A request carrying a tool allow-list is **refused** with
`UnsupportedToolRestriction`, because `codex exec` has no such flag and running
it with unrestricted tools would grant more authority than the caller asked
for. Nothing is spawned.
- `baikai-openai`: an evidence record's `thinking` field now describes what the
caller's reasoning-effort preference became on the wire for the specific host
the call went to, across **all seven** OpenAI-compatible wire shapes. The
OpenAI-native shape sends the canonical level verbatim and records no
adjustment, because it expresses every level exactly. The four shapes that
carry an effort word for a non-native host record `effort_clamped` whenever
the word differs from the canonical name — `minimal` becomes `low`, and both
`xhigh` and `max` become `high`. Z.ai and Qwen accept a bare
`enable_thinking: true` with no depth, so **every** level records
`effort_collapsed_to_toggle`: a caller asking for `max` and a caller asking
for `low` produce byte-identical requests there, and only the evidence record
can tell them apart. A host with no reasoning controls records
`thinking_dropped_unsupported_host` where the option previously vanished with
no trace. A forty-two-row table test pins the translation and the shaped
request body for every shape at every level.
- `baikai-openai`: the Chat Completions provider now fills in the evidence record
it previously left blank. It records the model **the host reported running**
(read from the first streamed chunk carrying a top-level `model` field and
never overwritten by a later one), the host's `x-request-id` correlation
header, the response id, the token counts the host actually reported, and a
commitment digest over the assembled response. A field the host did not report
stays `"unobserved"` and is never backfilled from the request — in particular,
a call that fails before any chunk arrives reports no observed model at all.
`strength` is `model_observed` when both the model and a correlation
identifier arrived, `correlated` when only the identifier did, and
`requested_only` otherwise; a 2xx status never raises it, because a 200 means
the request was accepted, not that any particular model ran.
`fully_observed` is unreachable on this transport, since no host in this
ecosystem echoes the reasoning configuration it applied.
- `baikai-openai`: new exports from `Baikai.Provider.OpenAI.Sse` —
`ResponseMetadata` and `capturedHeaderNames` — and from
`Baikai.Provider.OpenAI.Api` — `openaiChatStreamWith` and `SseDriver`.
Response-header capture is an **allow-list** (`x-request-id`, `request-id`,
`x-amzn-requestid`, `x-ms-request-id`, `cf-ray`, in that preference order),
not a denylist, so a header a future gateway adds is not recorded by default.
The list is longer than the Anthropic one because this transport speaks to an
open-ended set of hosts and the gateways commonly in front of them.
- `baikai-openai`: the same field for `codex exec`: mode `flag`, wire field
`model_reasoning_effort`, and **no** adjustments at any level. Codex is the
only transport in baikai that expresses all six canonical levels exactly, and
a test asserts each one reaches the command line verbatim.
### Fixed
- `baikai-openai`: `Response.responseId` was always `Nothing` on the Chat
Completions transport, although every compatible host sends a top-level `id`
on every streamed chunk. It now carries the identifier the host reported, on
both the successful and the failed terminal.
### Changed
- **Breaking:** `baikai-openai`: `Baikai.Provider.OpenAI.Sse`'s four streaming
entry points — `openaiSseStream`, `openaiSseStreamValue`,
`openaiSseStreamValueWithHeaders`, and `sseFromResponse` — take a new
`ResponseMetadata -> IO ()` callback immediately before the existing per-chunk
callback. It fires exactly once, before the first chunk, on both the success
and the non-2xx path — a failed call's correlation identifier is if anything
more valuable than a successful one's. Pass `(\_ -> pure ())` to keep the
previous behaviour. The callback is separate rather than a widening of the
per-chunk one because that one runs once per SSE frame and response-level data
does not belong on that path.
- **Breaking:** `baikai-openai`: `Baikai.Provider.OpenAI.Api`'s `RawChunk` gains
`model` and `responseId` fields, both `Maybe Text`. Code that pattern-matches
on `RawChunk` is unaffected; code that constructs one with record syntax must
add them.
- **Breaking:** `baikai-openai`: `Baikai.Provider.OpenAI.Shape`'s
`shapeRequestBody`, `streamRequestBody`, and `injectThinkingShape` now return
`(Aeson.Value, ThinkingTranslation)` instead of a bare body. Take `fst` to
keep the previous value. The description has to travel out of the shaping step
because nothing downstream can recompute it: it depends on the host's
`ThinkingFormat`, which only the compat lookup knows. **No request body
changed** — every one of the seven shapes puts exactly the same bytes on the
wire as before.
- **Breaking:** `baikai-openai`: `codexInteractiveCommand` now returns
`Either AgentRenderError (FilePath, [String])` and `launchCodexInteractive`
returns `IO (Either AgentRenderError InteractiveLaunchResult)`. A request
whose `safety` is a non-empty `ClaudeAllowedTools` list — which `codex` has
no flag for — is refused with `SafetyNotExpressible AgentCodex`, quoting the
rejected tools and suggesting `CodexSandbox` or `DefaultSafety`. Previously
the allow-list was silently discarded and Codex was started with its default
sandbox. The same `Left`/`Right` reading applies, `DefaultSafety` and an
empty allow-list are never refused, and no previously rendered argument
vector changed. Callers must handle the refusal branch.
Both changes make the interactive surface honor the same contract as the new
unattended surface: a safety policy the chosen provider cannot express fails
visibly instead of silently becoming a weaker policy. Downstream consumers
must adapt before upgrading; the known one is `shinzui/seihou`, whose
`Seihou.CLI.AgentLaunchExec` module builds interactive launch requests.
## [baikai-trace-otel 0.3.0.3] - 2026-08-05
### Added
- `baikai-trace-otel`: the sink attaches an evidence record's salient fields to
the open span as flat attributes (`baikai.evidence.run_id`,
`baikai.evidence.call_id`, `baikai.evidence.strength`, the two digests, and
`gen_ai.response.model` only when the provider actually reported one) rather
than serialising the record into one blob. A `CallEvidence` event neither
opens nor closes a span.
### Changed
- `baikai-trace-otel`: widened its `baikai` bound to admit `0.5`. No API change.
## [baikai-effectful 0.3.0.3] - 2026-08-05
### Changed
- Widened its `baikai` bound to admit `0.5`. No API change; the package's
own surface is untouched.
## [baikai-kit 0.1.0.4] - 2026-08-05
### Changed
- Widened its `baikai` bound to admit `0.5`. No API change; the package's
own surface is untouched.
## [baikai-agent 0.1.0.0] - 2026-08-05
### Added
- `baikai-agent`: **new package** (`0.1.0.0`) holding the unattended
coding-agent runner. `Baikai.Agent.Run.runAgentCommand` takes an
`AgentRunRequest` and an already-rendered `AgentCommand` and spawns the tool
with no terminal and no human present. It delivers the prompt on standard
input and closes the handle, drains standard output and standard error
concurrently so a chatty agent cannot deadlock on a full pipe, retains at most
`outputLimit` bytes per stream while reading and discarding the excess, and
honors the three output disciplines. Preconditions run before any spawn: a
missing working directory is `WorkingDirMissing` and unset or empty declared
variables are `MissingEnvironment`, listing all of them at once. On timeout
the child's whole process group is interrupted, given a grace period, and then
terminated, so the agent's own child processes go with it; the failure reports
the configured limit. A non-zero exit code is a successful run carrying that
code, not a failure. The runner consumes an already-rendered `AgentCommand`
and never imports a vendor renderer, so it is exercised entirely with
hand-written argument vectors. Its POSIX-signal escalation is conditional on a
non-Windows build.
- `baikai-agent`: new exposed module `Baikai.Agent.Config`, the layered
configuration layer. `resolveAgentJob` resolves one named job across five
layers — built-in defaults, the operator file, the repository file, the
environment, then command-line overrides, later layers winning — and returns
the resolved `AgentJob` together with a report attributing every value to the
file, line, and column it came from. `agentJobRequest` converts a job into an
`AgentRunRequest`, taking the prompt at call time. `listAgentJobs` enumerates
configured job names, sorted, each attributed to the highest-precedence scope
defining it. `defaultAgentConfigPaths` locates
`$XDG_CONFIG_HOME/baikai/agents.kdl` (or `$HOME/.config/baikai/agents.kdl`)
and `./.baikai/agents.kdl`, with no upward search through parent directories.
The **policy ceiling** is loaded by a separate function, `loadAgentCeiling`,
against a separate source list containing the operator file and nothing else:
no repository file, environment variable, or command-line override can raise
it. `applyCeilingToJob` refuses an over-broad request with `CeilingRejected`
rather than clamping it. With no operator file the ceiling is
`defaultAgentCeiling`. `safety.provider-args` is classified secret and renders
as `<redacted>` in any report or structured error.
New dependencies: `settei`, `settei-env`, `settei-kdl`, and
`settei-optparse-applicative` (all `^>=0.2`, published on Hackage at
`0.2.0.0`), plus `containers` and `filepath`. `settei-formats` is deliberately
excluded, because it bundles Dhall loading and repository configuration is
untrusted input here.
- `baikai-agent`: the **`baikai` executable**, with the `agent run`,
`agent show`, and `agent list` commands, and the `Baikai.Agent.Cli` module
that implements them. A shell script now invokes one stable command, supplies
a prompt on standard input, and selects Claude Code or Codex entirely through
configuration.
`agent run` resolves the named job, caps it against the operator ceiling,
renders it through the vendor renderer for its provider, and spawns it. The
agent's own exit code passes through unchanged; Baikai's own failures use 64
and above following the `sysexits` convention — 64 for a usage error or an
empty prompt, 69 when the executable could not be started, 70 for malformed
output, 75 for a timeout, 77 for a policy refusal, and 78 for a configuration
problem. The prompt comes from `--prompt-stdin`, `--prompt-file`, or
`--prompt`, which are mutually exclusive, and is decoded as UTF-8 explicitly
rather than through the handle's locale encoding.
`agent show` performs the whole pipeline except spawning and prints each
resolved value with the file, line, and column it came from, the policy
ceiling in force and where it was read, and the exact argument vector that
would be spawned — with `<redacted>` in place of any raw provider argument. A
job whose policy is refused prints its configuration first and then the
refusal. `agent list` enumerates configured jobs and the scope each came from.
Every Baikai diagnostic goes to standard error. The agent's own output follows
the job's output mode, so `response=$(baikai agent run job)` yields the
agent's answer alone for a capturing job. `--set KEY=VALUE` overrides one
setting of the selected job through `settei`'s own command-line source, so an
override is attributed with the same fidelity as a file. `--json` emits
exactly one JSON object per command.
New dependencies for `baikai-agent`: `baikai-claude`, `baikai-openai`, and
`optparse-applicative`. The provider packages are needed only so that
`renderJobCommand`, the single provider dispatch point in the codebase, can
reach both renderers. This is the first dependency in the workspace from
`baikai-agent` onto the provider packages, so `baikai-agent` now publishes
after all three of `baikai`, `baikai-claude`, and `baikai-openai`.
The user guide `docs/user/unattended-agent-runs.md` documents the whole
surface: the three commands with their flags, exit codes, and stream
discipline; the KDL job format and layer precedence; the operator ceiling and
redaction; the capability mapping tables for both tools; and a before-and-after
migration of a script that embeds provider flags today.
`docs/user/cli-providers.md` and `docs/user/interactive-launches.md` link to
it, and the capability mapping tables moved there from the latter.
- `baikai-agent`: **an unattended coding-agent run now produces model-call
evidence.** This surface previously had no observability of any kind: no trace
sink, no `Response`, no usage, no identifiers. An operator could show that a
process started, exited, and took some time; they could not show which model
ran, which reasoning effort was applied, or which agent session the run
corresponds to in the vendor's records.
A record carries the run and call identifiers, the resolved executable and its
own reported version, digests over the request, the requested model and what
the reasoning-effort request became on the command line, whatever the tool
reported about itself, the outcome, and an honest strength.
**A zero exit status never raises the strength.** On this surface that rule
matters more than anywhere else, because almost every unattended run exits
zero. A coding agent that exits zero has demonstrated that it ran, not which
model served it.
Two things gate what a record can prove, and neither is the default. The job
must **capture** output — under `inherit` the agent's bytes went to the
operator's terminal and baikai never held them — and the tool must be
configured to print a structured format, which means `--output-format json`
for `claude` or `--json` for `codex exec` through the job's `provider-args`.
Without both, the tool's session identifier, model, and token counts are
genuinely unavailable and the record says `"unobserved"` rather than inferring
anything. A timed-out run records `aborted`; a run that never started records
nothing at all.
- **Breaking:** `baikai-agent`: `Baikai.Agent.Run.runAgentCommand` takes two new
leading arguments and returns the new outcome type:
`Maybe EvidenceRequest -> ThinkingTranslation -> AgentRunRequest -> AgentCommand -> IO AgentRunOutcome`.
A caller who wants the previous behaviour passes `Nothing` and
`Baikai.Evidence.noThinkingRequested` and reads the `outcome` field; that path
is byte-for-byte what it was, and costs what it cost — no digest is computed,
no call identifier is generated, and the tool is not invoked a second time to
read its version.
- `baikai-agent`: `baikai agent run` gains `--evidence-file PATH` and
`--run-id TEXT`. Supplying neither leaves the run on the pre-existing path at
the pre-existing cost; supplying either turns recording on, with the job's own
name standing in as the run identifier when only a destination is given. The
file is written atomically — a staging file beside the destination, then a
rename — so a reader polling the path never sees a half-written object, and it
is never appended to. A failed write is reported on standard error and never
changes the exit code, because the agent's own status is what a calling script
branches on. `docs/user/unattended-agent-runs.md` documents both options and,
more importantly, what the record does and does not prove.
- `baikai-agent`: `baikai agent run` gains `--require-evidence STRENGTH`, taking
`requested_only`, `correlated`, `model_observed`, or `fully_observed` — the
same words a record's `strength` field spells, so what one record showed can
be passed back as the next run's requirement. A job whose configuration cannot
produce evidence of at least that strength is refused before anything is
spawned, exiting 77 — the code a ceiling violation and an inexpressible safety
policy already use, so a script branching on 77 needs no new case.
## [baikai-claude 0.4.0.1] - 2026-07-30
### Fixed
- Widened the `crypton` bound from `^>=1.0` to `>=1.0 && <1.2` so consumers can
build `baikai-claude` alongside packages that require `crypton` 1.1.x (for
example `pg-migrate-1.1.0.0`), which previously had no solvable build plan.
The only `crypton` use is `Crypto.Hash` (`Digest`, `SHA256`) in
`Baikai.Provider.Claude.Transport`, whose API is identical across the 1.0/1.1
boundary. No API change.
## [baikai 0.4.1.0] - 2026-07-20
### Changed
- Version bump only; no library API or code changes. Released so the umbrella
release tag `baikai-0.4.1.0` names a fresh core version alongside the breaking
`baikai-claude` / `baikai-openai` 0.4.0.0 releases, matching the tag
convention downstream consumers pin against.
## [baikai-claude 0.4.0.0] - 2026-07-20
### Changed
- **Breaking:** `claudeCliCommand` now takes the `Options` record and forwards
`Options.thinking` to batch `claude -p` as `--effort <level>` (`minimal`
collapses to `low`, matching the interactive launcher and the claude CLI's
lack of a `minimal` value). `thinking = Nothing` emits no effort flag, keeping
existing argv byte-for-byte. The added parameter is a PVP-major signature
change.
## [baikai-openai 0.4.0.0] - 2026-07-20
### Changed
- **Breaking:** `codexCliCommand` now takes the `Options` record and forwards
`Options.thinking` to `codex exec` as `-c model_reasoning_effort=<level>` for
all six effort levels. `thinking = Nothing` emits no override, keeping
existing argv byte-for-byte. The added parameter is a PVP-major signature
change.
## [baikai 0.4.0.0] - 2026-07-20
### Added
- Added `ThinkingXHigh` and `ThinkingMax` to the exported `ThinkingLevel`
vocabulary and added a defaulted `InteractiveLaunchRequest.effort` field.
Extending the closed sum type is a PVP-major API change for downstream
exhaustive matches.
## [baikai-claude 0.3.0.2] - 2026-07-20
### Added
- Added `--effort` rendering to interactive Claude Code launches and preserved
`xhigh` / `max` on native adaptive Anthropic API requests, with larger fixed
budgets for manual-thinking models.
### Changed
- Bumped the internal `baikai` dependency bound to `^>=0.4.0` for the
baikai 0.4.0.0 release.
## [baikai-openai 0.3.0.2] - 2026-07-20
### Added
- Added `model_reasoning_effort` overrides to interactive Codex launches and
preserved `xhigh` / `max` in native OpenAI request JSON; non-native
OpenAI-compatible request shapes continue to clamp them to `high`.
### Changed
- Bumped the internal `baikai` dependency bound to `^>=0.4.0` for the
baikai 0.4.0.0 release.
## [baikai-trace-otel 0.3.0.2] - 2026-07-20
### Changed
- Bumped the internal `baikai` dependency bound to `^>=0.4.0` for the
baikai 0.4.0.0 release. No API changes.
## [baikai-effectful 0.3.0.2] - 2026-07-20
### Changed
- Bumped the internal `baikai` dependency bound to `^>=0.4.0` for the
baikai 0.4.0.0 release. No API changes.
## [baikai-kit 0.1.0.3] - 2026-07-20
### Changed
- Bumped the internal `baikai` dependency bound to `^>=0.4.0` for the
baikai 0.4.0.0 release. No API changes.
## [baikai 0.3.1.0] - 2026-07-15
### Added
- Added `claude-sonnet-5` to the Anthropic model catalog (1M context window,
128k max output, `tool_call` + reasoning).
- Added the `gpt-5.6` family — `gpt-5.6`, `gpt-5.6-luna`, `gpt-5.6-sol`, and
`gpt-5.6-terra` — to the OpenAI model catalog (chat-completions with
`tool_call` support).
### Changed
- Corrected `claude-sonnet-4-5` context window to 1M tokens and
`claude-sonnet-4-6` max output to 128k tokens in the catalog.
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai-claude 0.3.0.1] - 2026-07-15
### Changed
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai-openai 0.3.0.1] - 2026-07-15
### Changed
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai-trace-otel 0.3.0.1] - 2026-07-15
### Changed
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai-effectful 0.3.0.1] - 2026-07-15
### Changed
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai-kit 0.1.0.2] - 2026-07-15
### Changed
- Added PVP-compliant upper bounds to all previously-unbounded library and
executable dependencies.
## [baikai 0.3.0.0] - 2026-07-03
### Added
- Added the documented record-update bases `emptyOptions`, `emptyContext`,
`emptyModel`, `emptyResponse`, `emptyTool`, `emptyTextContent`,
`emptyThinkingContent`, `emptyToolCall`, `emptyImageContent`,
`emptyEmbeddingModel`, plus zero-valued bases `zeroUsage`, `zeroCost`,
`zeroCostBreakdown`, and `zeroModelCost`.
- Added `firstEmbedding`, a total accessor for OpenAI-compatible embedding
responses.
- Added `responseError`, `errorResponse`, `httpError`, and
`parseRetryAfterSeconds` for the in-band error contract.
### Changed
- **Breaking:** Constructors for evolvable records are no longer exported:
`Options`, `Context`, `Model`, `OpenAICompletionsCompat`,
`AnthropicMessagesCompat`, and `InteractiveLaunchRequest` are built from
exported base values plus record updates.
- **Breaking:** The `_X` base values are deprecated in favor of the new
`empty*` and `zero*` names; the aliases remain for this release.
- **Breaking:** Removed `unModel`; use `mkModel` or `emptyModel` record
updates.
- **Breaking:** Renamed `InteractiveLaunchRequest.model` to `modelId`.
- **Breaking:** `Response.latencyMs` and trace event `latencyMs` fields are
now `Int`.
- **Breaking:** `completeRequest` / `completeRequestWith` no longer throw
`BaikaiError` for unregistered API tags; they return an error-shaped
`Response`.
- **Breaking:** CLI providers now report subprocess/decode/provider failures
in-band as error-shaped `Response`s.
- **Breaking:** `errorTerminal` now requires a `BaikaiError`, enforcing
structured error details for `EventError` construction sites.
- Documented that `Baikai.Prelude` is a convenience module outside the PVP
stability contract and that `.Internal` modules have no compatibility
guarantees.
### Fixed
- Empty embedding `data` arrays now produce a typed `decodeError` instead of
crashing on an empty vector.
- The model-fetch JSON renderer now delegates string escaping to aeson.
- The model generator now fails on sanitized Haskell identifier collisions
instead of rendering duplicate bindings.
- Live HTTP status, `Retry-After`, and network-failure classification now
works on both API providers.
- `content_filter` / Anthropic refusals terminate as classified `EventError`
terminals, and `liftCompleteToStream` preserves error-shaped responses.
## [baikai-claude 0.3.0.0] - 2026-07-03
### Changed
- **Breaking:** `Baikai.Provider.Claude.ErrorClass` moved to
`Baikai.Provider.Claude.Internal.ErrorClass`.
- **Breaking:** `mapRequest` and pure request-shaping helpers moved from
`Baikai.Provider.Claude.Api` to
`Baikai.Provider.Claude.Internal.Request`.
- **Breaking:** `ClaudeCliConfig` and `ClaudeInteractiveConfig` constructors
are no longer exported; start from their default config values and update
fields.
- **Breaking:** CLI and interactive `extraArgs` fields are now `[Text]`.
## [baikai-openai 0.3.0.0] - 2026-07-03
### Changed
- **Breaking:** `Baikai.Provider.OpenAI.ErrorClass` moved to
`Baikai.Provider.OpenAI.Internal.ErrorClass`.
- **Breaking:** `mapRequest` and pure request-shaping helpers moved from
`Baikai.Provider.OpenAI.Api` to
`Baikai.Provider.OpenAI.Internal.Request`.
- **Breaking:** `CodexCliConfig` and `CodexInteractiveConfig` constructors are
no longer exported; start from their default config values and update fields.
- **Breaking:** CLI and interactive `extraArgs` fields are now `[Text]`.
## [baikai-trace-otel 0.3.0.0] - 2026-07-03
### Changed
- Updated the `baikai` dependency bound to `^>=0.3.0`.
- Adjusted to the core trace event `latencyMs :: Int` type.
## [baikai-effectful 0.3.0.0] - 2026-07-03
### Changed
- Updated the `baikai` dependency bound to `^>=0.3.0`.
## [baikai-kit 0.1.0.1] - 2026-07-03
### Changed
- Updated the `baikai` dependency bound to `^>=0.3.0`.
## [baikai 0.2.0.0] - 2026-06-21
### Added
- `Usage`, `Cost`, and `CostBreakdown` now have `Semigroup`/`Monoid`
instances that add field-by-field, plus `sumUsage :: Foldable f => f
Usage -> Usage`, so callers can total per-call usage and cost.
`reasoningTokens` combines as presence-wins (`Nothing` only when both
operands are `Nothing`).
- A categorised error model: `BaikaiError` is now a record carrying an
`ErrorCategory` (`AuthError`, `RateLimited`, `ContextOverflow`,
`InvalidRequest`, `TransientError`, `DecodeFailure`, `ProcessFailure`,
`ProviderUnavailable`, `OtherError`), an optional HTTP `httpStatus`, a
`retryAfterSeconds` hint, and a subprocess `exitCode`. New smart
constructors (`providerError`, `invalidRequest`, `decodeError`,
`processError`, `rateLimited`, `authError`, `providerUnavailable`),
the `isRetryable` predicate, and the pure `classifyHttpStatus` /
`classifyHttpStatusWithBody` helpers let callers implement retry
policy without parsing error text. `ErrorCategory` and `BaikaiError`
serialize to JSON.
- `Response` and the streaming `EventError`'s `TerminalPayload` now
carry `errorInfo :: Maybe BaikaiError`, so a failed `completeRequest`
(or a drained stream) exposes the structured category/retry hint
in-band. `Baikai.Stream.Event` gains `doneTerminal` / `errorTerminal`
constructors.
### Changed
- **Breaking:** `BaikaiError`'s four flat constructors
(`ProviderError`, `RequestInvalid`, `DecodeError`, `ProcessError`)
were replaced by the record above. Migrate by lowercasing to the
smart constructors — `ProviderError "x"` becomes `providerError "x"`,
`ProcessError n "x"` becomes `processError n "x"`, etc.
- **Breaking:** `Baikai.Stream.Event.TerminalPayload` and
`Baikai.Response.Response` gained an `errorInfo` field; build
`TerminalPayload` via `doneTerminal` / `errorTerminal`.
### Fixed
- Restored JSON decoding for `BaikaiError` values with omitted optional
metadata fields.
## [baikai-claude 0.2.0.0] - 2026-06-21
### Added
- The Anthropic API and `claude -p` CLI providers now classify failures
into the typed `BaikaiError` categories: HTTP errors (via the caught
`servant-client` `ClientError`) map status/`Retry-After`/body onto
`AuthError` / `RateLimited` / `ContextOverflow` / `InvalidRequest` /
`TransientError`, and mid-stream Anthropic `error` events are
classified by their error type. The result is surfaced on
`Response.errorInfo`.
## [baikai-openai 0.2.0.0] - 2026-06-21
### Added
- The OpenAI/OpenAI-compatible API and `codex exec` CLI providers now
classify failures into the typed `BaikaiError` categories the same way
as `baikai-claude` (HTTP `ClientError` for status-based errors,
streamed error text for mid-stream errors), surfaced on
`Response.errorInfo`.
## [baikai-trace-otel 0.2.0.0] - 2026-06-21
### Changed
- Updated the `baikai` dependency bound to `^>=0.2.0` for compatibility with
the `baikai 0.2.0.0` breaking API release.
## [baikai-effectful 0.2.0.0] - 2026-06-21
### Changed
- Updated the `baikai` dependency bound to `^>=0.2.0` for compatibility with
the `baikai 0.2.0.0` breaking API release.
## [baikai 0.1.1.0] - 2026-06-12
### Added
- Added provider-agnostic `ResponseFormat` support on `Options`, including
plain JSON-object mode and named JSON-schema mode.
- Added `Baikai.Embedding`, an OpenAI `/v1/embeddings` client for text
embeddings.
## [baikai-claude 0.1.1.0] - 2026-06-12
### Added
- Mapped baikai `ResponseFormat` options onto Anthropic `output_config` for
Claude API requests.
- Exported `mapRequest` for request-mapping tests and downstream inspection.
## [baikai-openai 0.1.1.0] - 2026-06-12
### Added
- Mapped baikai `ResponseFormat` options onto OpenAI Chat Completions
`response_format`.
- Exported `mapRequest` for request-mapping tests and downstream inspection.
## [baikai-effectful 0.1.0.0] - 2026-06-12
### Added
- Initial release: effectful binding for baikai with the `Baikai` dynamic
effect, `complete`, `streamCollect`, `streamEach`, and registry-backed
interpreters.
## [baikai 0.1.0.0] - 2026-06-04
### Added
- Initial release: unified Haskell interface for working with multiple AI
providers. Core modules including `Baikai`, `Baikai.Prelude`, `Baikai.Api`,
`Baikai.Provider`, `Baikai.Provider.Registry`, `Baikai.Response`,
`Baikai.Stream`, `Baikai.Tool`, `Baikai.Trace`, and the cost/usage modules.
- Depends on released `streamly` (`>=0.11 && <0.13`) and `streamly-core`
(`>=0.3 && <0.5`) from Hackage, so all dependencies resolve from Hackage.
## [baikai-claude 0.1.0.0] - 2026-06-04
### Added
- Initial release: Anthropic Claude providers for the baikai abstraction,
wrapping the `claude` package for both the Anthropic API and the `claude -p`
CLI (`Baikai.Provider.Claude.Api`, `.Cli`, `.Interactive`).
## [baikai-openai 0.1.0.0] - 2026-06-04
### Added
- Initial release: OpenAI providers for the baikai abstraction, wrapping the
`openai` package for OpenAI's Chat Completions API
(`Baikai.Provider.OpenAI.Api`, `.Cli`, `.Interactive`).
## [baikai-trace-otel 0.1.0.0] - 2026-06-04
### Added
- Initial release: OpenTelemetry `TraceSink` adapter for baikai
(`Baikai.Trace.Sink.OpenTelemetry`), emitting one OTel span per provider call
with GenAI semantic-convention attributes plus baikai cost and latency.