Research Workbench
Research Workbench

Writing code

Having a model write code with tests, checked and retried until the tests pass, and running it safely in a sandbox.

Evidence and experimentspythoncsharpcodesandbox2 min read

Write code, on a project's header, asks a model to write code with tests. The workbench checks the code, and the model tries again with what the checks found, until the tests pass or the attempts run out. Every attempt, and every check, is kept with the project.

Asking

  1. What should the code do? For example: A function that tells whether a year is a leap year, following the Gregorian rules.
  2. Language: C# or Python. A language this computer cannot check says why.
  3. Attempts: 1 to 5 (3 by default).
  4. Model: the Writing code model in Settings, or another.
  5. Press Write it. It can take a minute or more, and you can follow it under Activity → Runs.

What is checked

Stage C# Python
Parses Roslyn the Python parser
Compiles the C# compiler byte-compiled
Analysers .NET code analysers pyflakes
Tests run in the sandbox pytest, in the sandbox

The result says It passed on the first attempt, It passed on attempt N, or It did not pass in N attempts. Each attempt shows each stage, what the tools said, and which tests failed. The code it ended with can be downloaded, or kept in the project's Generated code folder.

The sandbox

By default a model's code is only read and compiled, never run. To run its tests, tick Let the workbench run a model's code on this computer in Settings → This computer. The tests then run in a sandbox that cannot use the internet, open your files or see your API keys. It is stopped after 60 seconds or 512 MB of memory.

Running code needs Windows and the .NET SDK (and Python for Python code). The cloud workbench never runs a model's code.

To compare models or languages on the same tasks, write an experiment with the write-code workflow. See Experiments and runs.

Something went wrong. Reload the page to continue. Reload 🗙