packages feed

shikumi-optimize 0.3.0.0 → 0.3.0.1

raw patch · 19 files changed

+117/−102 lines, 19 filesdep ~baikaidep ~effectfuldep ~shikumi-compilePVP ok

version bump matches the API change (PVP)

Dependency ranges changed: baikai, effectful, shikumi-compile, shikumi-eval, shikumi-optimize, shikumi-trace

API changes (from Hackage documentation)

Files

CHANGELOG.md view
@@ -2,6 +2,10 @@  ## Unreleased +## 0.3.0.1 — 2026-10-05++- Move the dependency on `mori://shinzui/baikai/packages/baikai` to `>=0.7.1.0 && <0.8` and widen the `effectful` bound to `>=2.6 && <2.8`, so both effectful 2.6 and 2.7 are supported (effectful 2.7 needs `baikai-effectful` 0.4.0.2, effectful 2.6 needs 0.4.0.1). Bounds only; no source changed.+ ## 0.3.0.0 — 2026-09-08  - Raise the internal bounds to `shikumi ^>=0.4.0.0`, `shikumi-compile ^>=0.2.1.0`, `shikumi-eval ^>=0.3.0.0`, and `shikumi-trace ^>=0.3.0.0`.
shikumi-optimize.cabal view
@@ -1,8 +1,8 @@-cabal-version:   3.4-name:            shikumi-optimize-version:         0.3.0.0-synopsis:        The optimizer framework for shikumi LM programs (EP-10)-category:        AI+cabal-version: 3.4+name: shikumi-optimize+version: 0.3.0.1+synopsis: The optimizer framework for shikumi LM programs (EP-10)+category: AI description:   The optimizer framework for shikumi: search procedures that automatically   improve a 'Shikumi.Program.Program' by rewriting its per-node optimizable@@ -16,20 +16,26 @@   Every acceptance assertion runs fully offline against a deterministic stub LM,   so the whole package tests with no network and no API key. -license:         BSD-3-Clause-author:          Nadeem Bitar-maintainer:      nadeem@gmail.com-build-type:      Simple+license: BSD-3-Clause+author: Nadeem Bitar+maintainer: nadeem@gmail.com+build-type: Simple extra-doc-files: CHANGELOG.md  common common-options   ghc-options:-    -Wall -Wcompat -Widentities -Wincomplete-uni-patterns-    -Wincomplete-record-updates -Wredundant-constraints-    -fhide-source-paths -Wmissing-export-lists -Wpartial-fields+    -Wall+    -Wcompat+    -Widentities+    -Wincomplete-uni-patterns+    -Wincomplete-record-updates+    -Wredundant-constraints+    -fhide-source-paths+    -Wmissing-export-lists+    -Wpartial-fields     -Wmissing-deriving-strategies -  default-language:   GHC2024+  default-language: GHC2024   default-extensions:     DeriveAnyClass     DuplicateRecordFields@@ -37,8 +43,8 @@     OverloadedStrings  library-  import:          common-options-  hs-source-dirs:  src+  import: common-options+  hs-source-dirs: src   exposed-modules:     Shikumi.Optimize     Shikumi.Optimize.Bootstrap@@ -64,25 +70,28 @@     Shikumi.Optimize.Types    build-depends:-    , aeson            >=2.2      && <2.3-    , base             >=4.20     && <5-    , containers       >=0.6      && <0.9-    , effectful        >=2.5      && <2.7-    , generic-lens     >=2.2      && <2.4-    , lens             ^>=5.3-    , shikumi          ^>=0.4.0.0-    , shikumi-compile  ^>=0.2.1.0-    , shikumi-eval     ^>=0.3.0.0-    , shikumi-trace    ^>=0.3.0.0-    , text             ^>=2.1-    , vector           >=0.13     && <0.14+    aeson >=2.2 && <2.3,+    base >=4.20 && <5,+    containers >=0.6 && <0.9,+    effectful >=2.6 && <2.8,+    generic-lens >=2.2 && <2.4,+    lens ^>=5.3,+    shikumi ^>=0.4.0.0,+    shikumi-compile ^>=0.2.1.1,+    shikumi-eval ^>=0.3.0.1,+    shikumi-trace ^>=0.3.0.1,+    text ^>=2.1,+    vector >=0.13 && <0.14,  test-suite shikumi-optimize-test-  import:         common-options-  type:           exitcode-stdio-1.0+  import: common-options+  type: exitcode-stdio-1.0   hs-source-dirs: test-  main-is:        Main.hs-  ghc-options:    -threaded -with-rtsopts=-N+  main-is: Main.hs+  ghc-options:+    -threaded+    -with-rtsopts=-N+   other-modules:     AcceptanceSpec     BootstrapSpec@@ -106,19 +115,19 @@     StubLM    build-depends:-    , aeson-    , baikai            >=0.7.0.0  && <0.8-    , base-    , containers-    , effectful-    , generic-lens-    , lens-    , shikumi           ^>=0.4.0.0-    , shikumi-compile   ^>=0.2.1.0-    , shikumi-eval      ^>=0.3.0.0-    , shikumi-optimize  ^>=0.3.0.0-    , shikumi-trace     ^>=0.3.0.0-    , tasty-    , tasty-hunit-    , text-    , vector+    aeson,+    baikai >=0.7.1.0 && <0.8,+    base,+    containers,+    effectful,+    generic-lens,+    lens,+    shikumi ^>=0.4.0.0,+    shikumi-compile ^>=0.2.1.1,+    shikumi-eval ^>=0.3.0.1,+    shikumi-optimize ^>=0.3.0.1,+    shikumi-trace ^>=0.3.0.1,+    tasty,+    tasty-hunit,+    text,+    vector,
src/Shikumi/Optimize.hs view
@@ -1,7 +1,7 @@ -- | The public surface of the optimizer framework (EP-10). -- -- 'optimize' is the one stable entry point EP-12's CLI calls: it applies an--- 'Optimizer' strategy to a starting program and returns a 'CompiledProgram'. The+-- t'Optimizer' strategy to a starting program and returns a 'CompiledProgram'. The -- shared search-state plumbing ('selectBest', 'scoreOn', 'freezeProgram') lives in -- "Shikumi.Optimize.Search" and is re-exported here; the four strategies are -- re-exported from their own modules.
src/Shikumi/Optimize/Bootstrap.hs view
@@ -57,7 +57,7 @@ defaultBootstrapConfig = BootstrapConfig {passThreshold = 1.0, maxBootstrappedDemos = 4}  -- | Recover a demonstration from one teacher run: pair the typed input with the--- teacher's produced output, serialized to the JSON 'Demo' the run-time adapter+-- teacher's produced output, serialized to the JSON t'Demo' the run-time adapter -- decodes back into the node's typed demo channel. (The JSON keys are the record -- field names, so @fromModel@ round-trips them — see the unit test.) recoverDemo :: (ToJSON i, ToJSON o) => i -> o -> Demo@@ -71,7 +71,7 @@ -- stronger or chain-of-thought variant of the student, or the student itself; it -- must share the student's input/output types. Each teacher run reserves one -- predicted LM completion per teacher predict node before it runs; when the next--- teacher run does not fit the 'Budget', demo recovery stops and the demos found so+-- teacher run does not fit the t'Budget', demo recovery stops and the demos found so -- far are attached. bootstrapFewShotWith ::   (ToJSON i, ToJSON o) =>
src/Shikumi/Optimize/COPRO.hs view
@@ -8,7 +8,7 @@ -- -- COPRO consumes EP-19's grounded proposer ('Shikumi.Optimize.Propose.proposeInstructions') -- directly: each round's call passes the node's current instruction and its scored--- 'PastInstruction' history, and the proposer returns ranked candidates with the+-- t'PastInstruction' history, and the proposer returns ranked candidates with the -- current effective instruction always retained. Keeping that candidate writes no -- redundant override, preserving the safety property that a node never degrades. --@@ -42,7 +42,7 @@ import Shikumi.Optimize.Types (Budget (..), Optimizer (..), defaultBudget) import Shikumi.Program (Program, foldParams) --- | COPRO's two knobs plus the shared 'Budget'.+-- | COPRO's two knobs plus the shared t'Budget'. data CoproConfig = CoproConfig   { -- | candidate instructions generated per node per round (clamped to @>= 2@)     breadth :: !Int,@@ -60,7 +60,7 @@ -- | Coordinate-ascent instruction optimization. Visits each node in @foldParams@ -- order, optimizing it over @depth@ rounds against the already-improved earlier -- nodes. Proposer calls and candidate scoring reserve their predicted cost through--- one shared 'Budget', so the search returns the best-so-far before the next spend+-- one shared t'Budget', so the search returns the best-so-far before the next spend -- would exceed either ceiling. copro :: (ToJSON i, ToJSON o) => CoproConfig -> Optimizer i o copro cfg = Optimizer $ \train metric student -> do@@ -78,7 +78,7 @@ -- @breadth - 1@ fresh candidates (plus the retained current instruction) via the -- grounded proposer fed the scored attempt history, scores the not-yet-seen ones on -- the whole training set, records @(instruction, best-score)@, and sets the node to--- the best so far. Every spend is gated against the 'Budget'.+-- the best so far. Every spend is gated against the t'Budget'. optimizeNode ::   (ToJSON i, ToJSON o, LLM :> es, Concurrent :> es, Error ShikumiError :> es, Time :> es, Prim :> es) =>   CoproConfig ->
src/Shikumi/Optimize/GEPA.hs view
@@ -172,7 +172,7 @@  -- | The reflective evolutionary optimizer. Takes its reflective proposer and feedback -- metric explicitly (so it is testable under a stub LM) and returns V1's--- 'Optimizer'. GEPA gates its seed evaluation before any LM call; if the budget is+-- t'Optimizer'. GEPA gates its seed evaluation before any LM call; if the budget is -- too small to score the student once, it returns the student unscored. Each -- evolution step reserves a conservative full-step cost before capture, reflection, -- and child scoring.
src/Shikumi/Optimize/Instruction.hs view
@@ -6,7 +6,7 @@ -- kept. Optimization is greedy coordinate ascent — one node at a time, holding the -- others fixed — so the candidate count is linear in @nodes × proposals@. ----- The proposer and its signal-gatherers are themselves ordinary shikumi 'Program's, so+-- The proposer and its signal-gatherers are themselves ordinary shikumi t'Shikumi.Program.Program's, so -- they are typed, cached, traced, and testable with the same stub-LM machinery as -- everything else — the optimizer is written in the framework it optimizes. The -- /current/ effective instruction (override, or signature base when no override is@@ -18,12 +18,12 @@ -- completions per node (dataset summary, program describe, module describe, and one -- generation per proposal); scoring one candidate reserves one completion per dataset -- example per predict node. The search stops — returning the best found /so far/ —--- before either bound in the 'Budget' would be exceeded. When the remaining budget+-- before either bound in the t'Budget' would be exceeded. When the remaining budget -- cannot cover a node's full proposal, that node keeps its current instruction (no -- proposer call) rather than partially proposing. -- -- This module re-points V1's blind proposer at EP-19's grounded surface; the old--- @ProposeIn@/@ProposeOut@/@proposeInstruction@ predictor is removed (the grounded+-- @ProposeIn@\/@ProposeOut@\/@proposeInstruction@ predictor is removed (the grounded -- @GenerateInstructionIn@/@GenerateInstructionOut@ replaces it, still emitting a -- @proposedInstruction@ output field). module Shikumi.Optimize.Instruction
src/Shikumi/Optimize/KNN.hs view
@@ -2,7 +2,7 @@ -- examples whose inputs are most /semantically similar/ to it, and show those as the -- demos. Two forms: -----   * 'knnFewShot' — the faithful run-time form: a single 'Embed' node that, per+--   * 'knnFewShot' — the faithful run-time form: a single 'Shikumi.Program.Embed' node that, per --     input, embeds the input, ranks the training examples by cosine similarity, and --     runs the student under the @k@ nearest as demos. Demos depend on the input. --   * 'knnFewShotCentroid' — a compile-time fallback that bakes the @k@ examples@@ -10,7 +10,7 @@ --     for callers who cannot run an embedder at execution time. -- -- The embedder is injected as a /pure/ closure @Text -> Vector Double@ (the shape of--- EP-15's pure @runEmbedding@ argument), not the @Embedding@ effect: an 'Embed'+-- EP-15's pure @runEmbedding@ argument), not the @Embedding@ effect: an 'Shikumi.Program.Embed' -- body's row is fixed to @(LLM, Error ShikumiError)@, so it cannot call @embedText@; -- all embedding-effect work happens at the caller, outside the node. The run-time -- form carries no @Params@ (it serializes as an @Embed@ shape with an empty vector,@@ -87,7 +87,7 @@ centroid vs = V.map (/ fromIntegral (length vs)) (foldl1 (V.zipWith (+)) vs)  -- | The run-time KNN node: for each input, attach the @k@ nearest training examples--- as demos and run the student under them. A plain @Program i o@ (an 'Embed' node)+-- as demos and run the student under them. A plain @Program i o@ (an 'Shikumi.Program.Embed' node) -- usable anywhere a Program is; it carries no @Params@ (like @react@). knnDemos ::   (ToJSON i, ToJSON o, ToPrompt i) =>@@ -100,7 +100,7 @@   let exs = datasetExamples train    in embed $ \i -> runProgram (withDemos (nearestDemos embedder k exs (toPrompt i)) student) i --- | Run-time KNN as an 'Optimizer': selection is by embedding geometry, not by+-- | Run-time KNN as an t'Optimizer': selection is by embedding geometry, not by -- score, so it consults neither the metric nor the LM at optimize time and spends -- zero optimizer LM calls. The result is a structure-changing @Embed@ wrapper -- around the student. Its run-time selector closure is not persisted by
src/Shikumi/Optimize/LabeledFewShot.hs view
@@ -40,7 +40,7 @@     Just sc -> freezeProgram (withDemos (candidate sc) prog)  -- | The candidate demo sets a 'labeledFewShot' search considers: every size-@k@--- combination of the training examples (each turned into a JSON 'Demo'), in+-- combination of the training examples (each turned into a JSON t'Demo'), in -- deterministic enumeration order. Exposed so tests can reproduce the exact set of -- candidates the optimizer scored. labeledCandidateSets :: (ToJSON i, ToJSON o) => Int -> Dataset i o -> [[Demo]]
src/Shikumi/Optimize/MIPRO.hs view
@@ -20,12 +20,13 @@ -- integration point #4). -- -- __Search surrogate.__ DSPy drives phase 3 with Optuna's TPE sampler, which has no--- Haskell equivalent in this workspace. We implement __greedy coordinate descent with--- minibatch pruning__ over the joint grid: each trial screens the one-coordinate--- neighbours of the running best on a seeded minibatch, then full-evaluates the--- best-screened neighbour and accepts it only if it strictly improves the full score.+-- Haskell equivalent in this workspace. We implement+-- __greedy coordinate descent with minibatch pruning__ over the joint grid: each trial+-- screens the one-coordinate neighbours of the running best on a seeded minibatch,+-- then full-evaluates the best-screened neighbour and accepts it only if it strictly+-- improves the full score. -- This keeps the three essential, testable behaviours — joint grid, minibatch-screen--- + full-eval-confirm, and a hard 'Budget' — while staying fully deterministic. A+-- + full-eval-confirm, and a hard t'Budget' — while staying fully deterministic. A -- later EP can swap the neighbour-selection step for a TPE-lite without changing this -- module's public surface. module Shikumi.Optimize.MIPRO@@ -86,7 +87,7 @@ -- Configuration and presets -- --------------------------------------------------------------------------- --- | How aggressively to search. Mirrors DSPy's light/medium/heavy "auto" modes.+-- | How aggressively to search. Mirrors DSPy's light/medium\/heavy "auto" modes. data Miprov2Auto = Miprov2Light | Miprov2Medium | Miprov2Heavy   deriving stock (Eq, Show) @@ -109,7 +110,7 @@     maxBootstrappedDemos :: !Int,     -- | min metric score for a teacher run to contribute demos     bootstrapThreshold :: !Double,-    -- | hard predicted LM-completion / candidate ceiling (V1's 'Budget')+    -- | hard predicted LM-completion / candidate ceiling (V1's t'Budget')     budget :: !Budget   }   deriving stock (Eq, Show, Generic)@@ -240,7 +241,7 @@               }         pure cs --- | Render a recovered 'Demo' as @<input-json> => <output-json>@ for the proposal+-- | Render a recovered t'Demo' as @\<input-json\> => \<output-json\>@ for the proposal -- prompt's demo signal. renderDemo :: Demo -> Text renderDemo (Demo i o) = enc i <> " => " <> enc o@@ -256,7 +257,7 @@  -- | Search the joint per-node @(instruction × demoset)@ grid by greedy coordinate -- descent with minibatch screening, returning the best program found within the--- 'Budget'. Bootstrap teacher runs, grounded proposer calls, minibatch scoring, and+-- t'Budget'. Bootstrap teacher runs, grounded proposer calls, minibatch scoring, and -- full scoring all reserve predicted cost against one meter in 'miprov2With'. Each -- trial screens the one-coordinate neighbours of the running best on a seeded -- minibatch, then full-evaluates the best-screened neighbour and accepts it only if
src/Shikumi/Optimize/Pareto.hs view
@@ -1,5 +1,5 @@ -- | The Pareto-frontier bookkeeping for GEPA (EP-22): a pure, effect-free module so--- the frontier logic is trivially testable and reproducible. A 'Candidate' is a+-- the frontier logic is trivially testable and reproducible. A t'Candidate' is a -- program identified by its node-parameter vector, carrying its per-example score -- vector and aggregate. Keeping the /frontier/ (rather than one global best) -- preserves candidates that win on some examples even if not best on average — the
src/Shikumi/Optimize/Propose/Grounded.hs view
@@ -47,7 +47,7 @@ import Shikumi.Signature (mkSignature)  -- | The grounded proposer's input: every signal about the optimization target,--- rendered into the prompt under its field name by the generic 'ToPrompt'.+-- rendered into the prompt under its field name by the generic t'ToPrompt'. data GenerateInstructionIn = GenerateInstructionIn   { datasetDescription :: !Text,     programCode :: !Text,
src/Shikumi/Optimize/Propose/Summarize.hs view
@@ -1,12 +1,12 @@ {-# LANGUAGE FlexibleContexts #-}  -- | The signal-gatherers of the grounded proposer (EP-19) that are themselves typed--- Shikumi 'Program's — preserving V1's "the optimizer is written in the framework it+-- Shikumi t'Program's — preserving V1's "the optimizer is written in the framework it -- optimizes" pattern. Each mirrors a DSPy proposer sub-module: -- --   * 'renderProgramPseudo' / 'programDescriber' — DSPy's @DescribeProgram@: render --     the whole program as deterministic pseudo-code and describe what it does.---   * 'datasetDescriber' / 'observationSummarizer' / 'datasetSummary' — DSPy's+--   * 'datasetDescriber' \/ 'observationSummarizer' \/ 'datasetSummary' — DSPy's --     @DatasetDescriptor@ + @ObservationSummarizer@: observe patterns across sampled --     rows, then condense them into a 2-3 sentence summary. --   * 'moduleDescriber' — DSPy's @DescribeModule@: describe one node's role within@@ -68,7 +68,7 @@ -- ---------------------------------------------------------------------------  -- | Render a program as a short, deterministic, human-readable outline: one line--- per node and combinator, with 'Predict' nodes shown as @predict(inputs) ->+-- per node and combinator, with 'Shikumi.Program.Predict' nodes shown as @predict(inputs) -> -- outputs@ (using 'programFieldNames') and combinators shown by name. Shikumi's -- analogue of DSPy's @get_dspy_source_code@. Deterministic, so tests can assert on it. renderProgramPseudo :: Program i o -> Text@@ -95,11 +95,11 @@        in ((name <> ":") : map ("  " <>) lns, rest)     commas xs = if null xs then "?" else T.intercalate ", " xs --- | Render a single dataset example as @<input-json> => <expected-json>@.+-- | Render a single dataset example as @\<input-json\> => \<expected-json\>@. renderExampleRow :: (ToJSON i, ToJSON o) => Example i o -> Text renderExampleRow (Example i o) = encodeJsonText i <> " => " <> encodeJsonText o --- | Encode any 'ToJSON' value to compact 'Text'.+-- | Encode any t'ToJSON' value to compact 'Text'. encodeJsonText :: (ToJSON a) => a -> Text encodeJsonText = TL.toStrict . encodeToLazyText 
src/Shikumi/Optimize/Propose/Types.hs view
@@ -2,8 +2,8 @@ -- field-metadata accessor (integration point #3), the instruction-history vocabulary, -- and the proposer's request/result records. ----- This is the contract MIPROv2 (@docs/plans/20-miprov2-optimizer.md@) and COPRO--- (@docs/plans/21-copro-instruction-optimizer.md@) both consume, so it lives in its+-- This is the contract MIPROv2 (@docs\/plans\/20-miprov2-optimizer.md@) and COPRO+-- (@docs\/plans\/21-copro-instruction-optimizer.md@) both consume, so it lives in its -- own module — neither optimizer drags in @instructionSearch@'s loop to use it. module Shikumi.Optimize.Propose.Types   ( -- * Per-node field metadata (integration point #3)@@ -26,8 +26,8 @@ import GHC.Generics (Generic) import Shikumi.Program (NodeFields (NodeFields), Program, nodeFieldsIndexed) --- | A node's input/output field names, recovered structurally. A 'Predict' node--- hides its @i@/@o@ types existentially, so this carries the field /names/ (plain+-- | A node's input/output field names, recovered structurally. A 'Shikumi.Program.Predict' node+-- hides its @i@\/@o@ types existentially, so this carries the field /names/ (plain -- 'Text'), never a typed @Signature@ — exactly what EP-16's @nodeFieldsIndexed@ -- returns. data NodeFieldNames = NodeFieldNames@@ -36,7 +36,7 @@   }   deriving stock (Eq, Show, Generic) --- | One 'NodeFieldNames' per 'Predict' node, in @foldParams@/@mapParamsAt@ order+-- | One t'NodeFieldNames' per 'Shikumi.Program.Predict' node, in @foldParams@/@mapParamsAt@ order -- (integration point #3). Delegates to EP-16's @nodeFieldsIndexed@; the count and -- ordering align with @foldParams@ by construction, so -- @programFieldNames prog !! k@ describes the node @mapParamsAt k@ edits.@@ -62,7 +62,7 @@   }   deriving stock (Eq, Show, Generic) --- | Render the instruction history as lines @"score 0.83 :: <instruction>"@, capped+-- | Render the instruction history as lines @"score 0.83 :: \<instruction\>"@, capped -- at @maxInHistory@ entries. Empty history renders as @"No previous instructions."@ -- so the proposal prompt is always well-formed. renderHistory :: Int -> [PastInstruction] -> Text
src/Shikumi/Optimize/RandomSearch.hs view
@@ -1,5 +1,5 @@ -- | Bootstrap few-shot with random search (EP-23, DSPy's--- @BootstrapFewShotWithRandomSearch@): run V1's 'bootstrapFewShot' several times with+-- @BootstrapFewShotWithRandomSearch@): run V1's 'Shikumi.Optimize.Bootstrap.bootstrapFewShot' several times with -- different deterministic seeds — each shuffling the trainset and picking a random -- demo count — score each resulting program, and keep the best. A zero-shot baseline -- candidate is always included, so the search can never do worse than zero-shot.@@ -65,7 +65,7 @@         (x : _) -> lo + (x `mod` span')         [] -> lo --- | 'bootstrapRandomSearch' with explicit tunables. One 'Budget' covers all seed+-- | 'bootstrapRandomSearch' with explicit tunables. One t'Budget' covers all seed -- bootstrap teacher runs and the final candidate scoring pass; when the meter is -- exhausted, later seeds or scoring candidates are skipped and the best scored -- candidate so far is returned.@@ -93,7 +93,7 @@     Just sc -> freezeProgram (candidate sc)  -- | Run V1 bootstrap over @numCandidates@ random seeds plus a zero-shot baseline,--- score each on the dataset, and keep the best-scoring 'CompiledProgram'.+-- score each on the dataset, and keep the best-scoring t'Shikumi.Compile.Types.CompiledProgram'. bootstrapRandomSearch ::   (ToJSON i, ToJSON o) =>   Program i o ->
src/Shikumi/Optimize/Search.hs view
@@ -1,7 +1,7 @@ -- | The shared search-state plumbing every concrete optimizer reuses: the pure -- 'selectBest' fold, the @evaluate@-backed 'scoreOn' scorer, and 'freezeProgram'. ----- This module sits /below/ both 'Shikumi.Optimize' (which re-exports it) and the+-- This module sits /below/ both "Shikumi.Optimize" (which re-exports it) and the -- four optimizer modules (which import it), so there is no import cycle: the -- optimizers depend on the plumbing, not on the public driver. --@@ -92,7 +92,7 @@       else (calls, False)  -- | The predicted LM-call cost of scoring @p@ over @ds@ once: one call per--- example per 'Predict' node. Wrappers that re-run the LM can spend more.+-- example per 'Shikumi.Program.Predict' node. Wrappers that re-run the LM can spend more. scoringCost :: Dataset i o -> Program i o -> Int scoringCost ds p = datasetSize ds * max 1 (length (foldParams p)) 
src/Shikumi/Optimize/Types.hs view
@@ -1,12 +1,12 @@ {-# LANGUAGE RankNTypes #-} --- | The central abstractions of the optimizer framework (EP-10): the 'Optimizer'--- strategy object, the search 'Budget', and the 'Scored' candidate wrapper.+-- | The central abstractions of the optimizer framework (EP-10): the t'Optimizer'+-- strategy object, the search t'Budget', and the t'Scored' candidate wrapper. ----- An 'Optimizer' is a /search procedure/: given a training 'Dataset', a 'Metric',--- and a starting 'Program', it proposes new node parameters (instructions and+-- An t'Optimizer' is a /search procedure/: given a training 'Dataset', a 'Metric',+-- and a starting t'Program', it proposes new node parameters (instructions and -- few-shot demonstrations), scores each candidate by running the program over the--- dataset, and returns the best-scoring 'CompiledProgram' it found. An optimizer+-- dataset, and returns the best-scoring t'CompiledProgram' it found. An optimizer -- normally changes parameters while preserving boundary types. -- -- The additive 'Shikumi.Optimize.Structure.structureSearchWith' API selects among@@ -37,9 +37,9 @@ -- /is/ a call to @evaluate@ (via 'Shikumi.Optimize.scoreOn'). This is wider than -- the plan's original @(LLM :> es)@ sketch: the delivered EP-8 runner threads -- 'Concurrent' (bounded parallelism), 'Time' (monotonic per-example latency via--- the 'Shikumi.Effect.Time' effect), and 'Prim' (the usage-accounting 'IORef' in--- 'Shikumi.Eval.Usage.withUsageTotals'). 'IOE' is no longer required — it is--- supplied only at the discharge edge by 'runEff' under 'runTime'/'runPrim'. See+-- the "Shikumi.Effect.Time" effect), and 'Prim' (the usage-accounting 'Data.IORef.IORef' in+-- 'Shikumi.Eval.Usage.withUsageTotals'). 'Effectful.IOE' is no longer required — it is+-- supplied only at the discharge edge by 'Effectful.runEff' under 'Shikumi.Effect.Time.runTime'/'Effectful.Prim.runPrim'. See -- the plan's Decision Log. module Shikumi.Optimize.Types   ( Optimizer (..),
test/Miprov2Spec.hs view
@@ -52,7 +52,7 @@ -- Fixtures and run helpers -- --------------------------------------------------------------------------- --- | Region A (good/bad) needs a RULE instruction; region B (great/terrible) needs a+-- | Region A (good\/bad) needs a RULE instruction; region B (great\/terrible) needs a -- covering demo. See "StubLM".'StubLM.runJointStubLM'. jointTrain :: Dataset Sentence Label jointTrain =
test/StubLM.hs view
@@ -3,9 +3,10 @@ -- | Shared, network-free fixtures and a deterministic stub @LLM@ interpreter for -- the EP-10 optimizer suite. ----- The whole point of an optimizer test is that __changing a program's parameters--- changes its score__ — otherwise an optimizer cannot demonstrably improve--- anything. So the stub is not a constant: it inspects the rendered request+-- The whole point of an optimizer test is that+-- __changing a program's parameters changes its score__ — otherwise an optimizer+-- cannot demonstrably improve anything. So the stub is not a constant: it inspects+-- the rendered request -- (system prompt + messages) and answers accordingly, by a rule that is monotone -- in parameter quality. The task is binary sentiment classification of a -- 'Sentence' into a 'Label' (@"positive"@ / @"negative"@). The ground-truth rule