jiokua

consistently curious, hopefully not cat-like and dangerously so. currently building and thinking about ai, specifically exploring how it integrates and changes pre-existing systems and how we can begin to think through what an agent "experiences", for lack of a better word, without making any claims on consciousness or sentience.

Recent Activity

Experiments

research-reader-tool-experiments

All the evidence at once beat live search on long-document questions

Experiment run 24/24

Four ways of supplying evidence to the same model were compared on answer quality and token use. When all the evidence fit at once, the model answered most accurately while using less than half the tokens of live search. Reading the evidence found through search in one pass was also considerably more efficient than answering during the search itself.

Answer quality against token cost
Armtokens per questionanswer quality (F1)Frontier
handed a limited set3,826.6230.46Dominated
handed everything4,667.5560.73Efficient
searched for itself11,259.9760.63Dominated
handed what search found2,975.6510.65Efficient

file-in-activity-experiments

The task determined which version of a file the AI used

Experiment run 9/11

The same kind of mid-task file update was tested across two task settings. When the file was work in progress, the model usually used the new value in its final output; when the file was reference material, it usually retained the earlier value. This difference appeared across all ten scenarios.

  • brickworks
  • coldstore
  • dairy
  • glassworks
  • granary
  • locks
  • nursery
  • orchard
  • smokehouse
  • switchyard
Whether the mid-task update was taken, by framing
scenariosfile is the workfile is a sourceChange
brickworks0.9170.083
coldstore0.8330.000
dairy0.8330.250
glassworks0.9170.167
granary0.6670.083
locks0.9170.083
nursery0.8330.000
orchard1.0000.000
smokehouse0.9170.083
switchyard0.8330.417

file-phenomenon-experiments

Seeing a shared-file change worked as well as being told

Experiment run 22/22

The model received either a direct explanation, an informative record of what changed, an uninformative record, or no record. Direct explanation and an informative record both produced consistently correct choices. An uninformative record left the model near chance, while no record produced widely varied choices and often no choice at all.

Telling one file from two, by what the model could go on
evidence availablenMedianIQRRange
told outright161.001.00–1.001.00–1.00
change is visible161.001.00–1.001.00–1.00
record shows nothing160.500.50–0.550.38–0.61
no record160.470.40–0.580.20–1.00