jiokua
consistently curious, hopefully not cat-like and dangerously so. currently building and thinking about ai, specifically exploring how it integrates and changes pre-existing systems and how we can begin to think through what an agent "experiences", for lack of a better word, without making any claims on consciousness or sentience.
Recent Activity
Ideas panel selected
Experiments
All the evidence at once beat live search on long-document questions
Experiment run 24/24
Four ways of supplying evidence to the same model were compared on answer quality and token use. When all the evidence fit at once, the model answered most accurately while using less than half the tokens of live search. Reading the evidence found through search in one pass was also considerably more efficient than answering during the search itself.
The task determined which version of a file the AI used
Experiment run 9/11
The same kind of mid-task file update was tested across two task settings. When the file was work in progress, the model usually used the new value in its final output; when the file was reference material, it usually retained the earlier value. This difference appeared across all ten scenarios.
Seeing a shared-file change worked as well as being told
Experiment run 22/22
The model received either a direct explanation, an informative record of what changed, an uninformative record, or no record. Direct explanation and an informative record both produced consistently correct choices. An uninformative record left the model near chance, while no record produced widely varied choices and often no choice at all.