jupyter_json_default
def jupyter_json_default(
obj
):JSON serializer fallback matching Jupyter Python object handling.
Jupyter’s wire protocol accepts Python values that the standard JSON encoder rejects. jupyter_json_default converts them as Jupyter does:
int or floatJSON serializer fallback matching Jupyter Python object handling.
A notebook is just a json file.
It contains two sections, the metadata…:
…and, more importantly, the cells:
[{'cell_type': 'markdown',
'id': '801558df',
'metadata': {},
'source': ['## A minimal notebook']},
{'cell_type': 'code',
'execution_count': None,
'id': 'e2147a69',
'metadata': {'time_run': '2026-01-04T20:52:49.901559+00:00'},
'outputs': [{'data': {'text/plain': ['2']},
'execution_count': 0,
'metadata': {},
'output_type': 'execute_result'}],
'source': ['# Do some arithmetic\n', '1+1']}]
The second cell is a code cell with no outputs, because it hasn’t been executed yet. Executing a notebook needs the format nbclient expects, where some dict keys are available as attributes and others as ordinary keys. nbformat usually does this conversion, but it’s slow and inflexible. Our own version builds on fastcore’s dict2obj, which makes every key available both as an attribute and as a key.
langs holds each notebook language’s comment characters, taken from Quarto’s table. nb_lang reads a notebook’s language from its kernelspec, defaulting to Python. Each cell records its language in lang_, which is the notebook’s language unless the cell’s own metadata.language overrides it. The directive functions below use lang_ to recognize comments.
dict subclass that also provides access to keys as attrs, and has a pretty markdown repr
We use an AttrDict subclass which has some basic functionality for accessing notebook cells.
Two cells with the same id are equal, even when their content differs. Because equality follows the id, list.index and list.remove, which Notebook’s cell-moving methods use, find the right cell even when another cell has identical source. Cells without an id compare by source and cell_type.
c1 = NbCell(0, dict(cell_type='code', source='a', id='x'))
c2 = NbCell(0, dict(cell_type='code', source='b', id='x'))
c3 = NbCell(0, dict(cell_type='code', source='a', id='y'))
test_eq(c1, c2) # same id, different content -- still equal
assert c1 != c3 # same content, different id -- not equal
cells = [c1, c3] # duplicate content ('a'), different ids
test_eq(cells.index(c3), 1) # finds the actual cell, not the first with matching contentConverting our JSON to this nbclient-compatible format gives cells that pretty-print their source code:
{ 'cell_type': 'code',
'execution_count': None,
'id': 'e2147a69',
'idx_': 1,
'lang_': 'python',
'metadata': {'time_run': '2026-01-04T20:52:49.901559+00:00'},
'outputs': [ { 'data': {'text/plain': '2'},
'execution_count': 0,
'metadata': {},
'output_type': 'execute_result'}],
'source': '# Do some arithmetic\n1+1'}The abstract syntax tree of source code cells is available in the parsed_ property:
This reads the JSON for the file at path and converts it with dict2nb. For instance:
"{'cell_type': 'markdown', 'id': '801558df', 'metadata': {}, 'source': '## A minimal notebook', 'idx_': 0, 'lang_': 'python'}"
The file name read is stored in path_:
Create an NbCell containing text
Returns an empty new notebook
nb_frontmatter merges a notebook’s frontmatter from three sources. From lowest precedence to highest, they are the notebook’s own metadata.nbdev mapping, its first markdown cell, and its first raw cell anywhere in the notebook. As in directives, where comments win over the metadata form, a cell’s content wins over metadata.
YAML parsing applies only to a cell that is entirely a literal --- block, never to prose that merely starts with ---. Malformed YAML in such a block raises an error. A markdown cell that isn’t a --- block contributes its # title, > description and - key: value lines instead. With strvals=True, every YAML scalar stays a string except true, True, false and False, which frontmatter still turns into booleans.
Metadata values pass through unchanged. Unlike a cell directive, a frontmatter value of 'true' stays 'true' and isn’t normalized to ''.
nb_frontmatter accepts anything whose cells items have cell_type and source, including aidialog and solveit dialogs, whose messages work as cells. A message without a metadata attribute contributes nothing from metadata. nbdev’s FrontmatterProc also uses the single-cell pieces, cell_frontmatter and md_frontmatter.
Frontmatter from nb, merged lowest-to-highest from its metadata.nbdev mapping (values verbatim), first markdown cell (literal --- block, or # title synthesis), and first raw cell anywhere
Frontmatter synthesized from an H1-formatted markdown cell: # title, > description, and - key: value lines
Frontmatter mapping from a cell source that is entirely a literal --- block, else {}
fm_nb = new_nb([mk_cell('---\ntitle: T\nformdata:\n who: Sam\n n: 1000\n---', 'raw'), mk_cell('Body', 'markdown')])
test_eq(nb_frontmatter(fm_nb), {'title':'T', 'formdata':{'who':'Sam', 'n':1000}})
test_eq(nb_frontmatter(fm_nb, strvals=True)['formdata'], {'who':'Sam', 'n':'1000'})
test_eq(nb_frontmatter(new_nb([])), {})
h1 = mk_cell('# My title\n\n> A subtitle\n\n- author: Zoë\n\nProse.', 'markdown')
test_eq(nb_frontmatter(new_nb([h1])), {'title':'My title', 'description':'A subtitle', 'author':'Zoë'})
test_eq(md_frontmatter('# T\n- n: 1000\n- f: true', strvals=True), dict(title='T', n='1000', f=True)) # strvals honored via the `- key: value` path too
both = new_nb([h1, mk_cell('code'), mk_cell('---\ntitle: Raw wins\n---', 'raw')])
test_eq(nb_frontmatter(both)['title'], 'Raw wins') # first raw anywhere; raw beats markdown
test_eq(nb_frontmatter(both)['description'], 'A subtitle') # markdown keys merge beneath
meta_nb = new_nb([mk_cell('---\neval: false\n---', 'raw')], meta={'nbdev': {'eval':'true', 'export':'true'}})
test_eq(nb_frontmatter(meta_nb), {'eval':False, 'export':'true'}) # metadata.nbdev merges lowest, values verbatim (no bare-'' form)
test_eq(nb_frontmatter(new_nb([mk_cell('---\nt: v\n---\ntrailing', 'raw')])), {}) # whole-cell anchoring: a body disqualifies
test_eq(nb_frontmatter(new_nb([mk_cell('---\nt: v\n---', 'markdown')])), {'t':'v'}) # a markdown cell that IS a literal block parses
test_fail(lambda: nb_frontmatter(new_nb([mk_cell('---\nbad: "unclosed\n---', 'raw')])))Use new_nb to create a notebook in memory, without first writing one to disk and reading it back.
nbdev and Quarto put directives in comments at the top of a cell, such as #| export or #| eval: false. The directives end at the first line that isn’t a comment. #| foo: bar, #| foo:bar and #| foo bar all parse to the same directive, whose value is the raw text after its name. Because a value of true means the same as no value, no directive can have the literal value true: #| hide, #| hide: true and #| hide: are one directive.
Directives can also be stored in cell metadata, as a dict under an nbdev key, with the same values. Every value there must be a string, with "true" meaning a bare directive. A non-string value, such as a JSON false, raises an error instead of being read as "false". When the same name appears in both places, the comment wins.
get first line number where code occurs, where code_list is a list of code
_directive parses one directive line into its name and value. Both #| export: utils and #| export utils give the value utils: the raw text after the name and optional colon, otherwise unchanged apart from one leading space and any trailing whitespace. A value of true becomes '', the same as a bare directive. Lines that aren’t directives, including cell magics, parse to None.
test_eq(_directive('#| export: utils'), ('export','utils'))
test_eq(_directive('#| export utils'), ('export','utils'))
test_eq(_directive('#| export:utils'), ('export','utils'))
test_eq(_directive('#| hide'), ('hide',''))
test_eq(_directive('#| hide:'), ('hide',''))
test_eq(_directive('#| hide: true'), ('hide',''))
test_eq(_directive(' # | woo:baz'), ('woo','baz'))
test_eq(_directive('#| fig-cap: "two spaces: kept"'), ('fig-cap','"two spaces: kept"'))
test_eq(_directive('#| filter_stream secret apikey'), ('filter_stream','secret apikey'))
test_eq(_directive('%%timeit'), None)
test_eq(_directive('# plain comment'), None)dir_tag renders meta-form directives as the compact bracket tag shown after the type character in summary rows, as in c[export]:.... CellRow below and aidialog’s message previews both use it.
Meta-form nbdev directives in meta as a compact [k k=v] bracket tag, or '' if none
'[export]'
directives is a plain dict from name to value string, such as {'export': 'utils', 'hide': ''}, where a bare directive has the value ''. To edit, modify the copy that the getter returns and assign it back. The setter rewrites the comment block in canonical colon form, leaving any cell magic first and unchanged. Directives that came from cell metadata, which the cell tracks by name, go back to metadata instead of to comments. Names added through the setter become comments.
Metadata directives merge with comment directives, with the comment winning on conflict. Each edit goes back to wherever its directive came from. In metadata, the setter writes a bare directive as "true":
c = mk_cell('#| export: utils\n1', metadata=dict(nbdev=dict(export='other', eval='false')))
test_eq(c.directives, {'export':'utils', 'eval':'false'})
d = c.directives
d['eval'] = ''
d['hide'] = ''
c.directives = d
test_eq(c.metadata['nbdev'], {'eval': 'true'})
test_eq(c.source, '#| export: utils\n#| hide\n1')
test_eq(c.directives, {'export':'utils', 'hide':'', 'eval':''})
with expect_fail(TypeError, 'must be str'): mk_cell('1', metadata=dict(nbdev=dict(eval=False))).directivesStrip directives from source, keeping cell magics; with quarto, instead materialize every directive (metadata included) as a Quarto option line
Value of directive name ('' if bare), or default if absent
directive returns a directive’s value and has_directive reports whether it’s present. Use has_directive to test for presence, because a bare directive’s value is the falsy ''.
remove_directives has two modes. By default it strips every directive line, as module export needs clean source. With quarto=True it instead rewrites the cell with every directive, from comments and metadata alike, as a Quarto option line, writing bare ones as name: true. Quarto uses the options it knows and turns unknown ones into data-* attributes without warning.
c = mk_cell('%%time\n#| exports: utils\n#| hide\n#| code-fold: show\nslow()', metadata=dict(nbdev=dict(echo='false')))
test_eq(c.directive('exports'), 'utils')
assert c.has_directive('hide') and not c.has_directive('eval')
c.remove_directives(quarto=True)
test_eq(c.source, '%%time\n#| exports: utils\n#| hide: true\n#| code-fold: show\n#| echo: false\nslow()')
c = mk_cell('%%time\n#| exports: utils\n#| hide\n#| code-fold: show\nslow()', metadata=dict(nbdev=dict(echo='false')))
c.remove_directives()
test_eq(c.source, '%%time\nslow()')This returns the exact same dict as is read from the notebook JSON.
To save a notebook we first need to convert it to a str:
This returns the exact same string as saved by Jupyter.
The cell tools apply fastcore.tools’ string editing primitives to one notebook cell’s source. They match that module’s file tools, with the same operations and parameters, but take path, cell_id where the file tools take path. Every editor, including the structural cell_ast_replace, returns a diff of its change. view_cell shows a cell’s source with optional line numbers or exhash addresses.
fastcore.editskill documents the naming and parameter conventions shared across the editing toolkit, and re-exports this module’s editing tools.
Every tool in this family finds its unit by id. It accepts an exact match or any unique prefix, and raises a KeyError naming the problem otherwise. find_id implements that rule, both for the notebook tools here and for the dialog tools built on them:
The item with id k; a missing or ambiguous id raises KeyError naming the problem
Wrap text editor f as a cell editing function: path, cell_id addressing, diff-or-error return
The cell editors have the same signatures as their fastcore.tools counterparts. In doc() overviews here, their six shared options appear as one replace_params group, as they do across the toolkit:
def view_cell(
path:str, # Notebook file to read (expands `~`)
cell_id:str, # Id of the cell to view (exact, or unique prefix)
start_line:int=1, # Starting line to view
end_line:int=None, # End line (defaults to last line if None; may be past EOF, which clamps to the last line)
nums:bool=True, # Show line numbers?
lnhashs:bool=False, # Show exhash `lineno|hash|` addresses instead of line numbers?
incl_out:bool=False, # Append the cell's outputs in an `<out>` block?
trunc_out:bool=True, # Truncate included outputs to ~512 chars?
):View a cell’s source, optionally limited to 1-based line range
The examples below edit a scratch notebook. view_cell shows one cell’s source by id, with plain line numbers by default, for a line range, or with lnhash addresses for cell_exhash:
tmp_nb = Path('tmp_cells.ipynb')
write_nb(new_nb([mk_cell('a=1\nprint(a)'), mk_cell('# title', 'markdown')]), tmp_nb)
cid = read_nb(tmp_nb).cells[0].id
test_eq(str(view_cell(tmp_nb, cid)), '1: a=1\n2: print(a)')
test_eq(str(view_cell(tmp_nb, cid, start_line=2, nums=False)), 'print(a)')
test_eq(str(view_cell(tmp_nb, cid, lnhashs=True)).splitlines()[0], lnhash(1,'a=1')+'a=1')Each editor works as one transaction: it applies the edit, writes the file and returns a diff. Check the edit in the returned diff, with no second read of the file:
The line-addressed editors follow the text primitives’ conventions. Inserting at line 0 prepends a line that deleting line 1 removes again. cell_replace_lines with no range replaces the whole source with a block of lines ending in a newline:
An unmatched str_replace returns an error instead of writing the file. A missing cell id raises a KeyError:
Here’s how to put all the pieces of fastcore.nbio together:
[{'cell_type': 'code', 'execution_count': None, 'id': 'd0c314de', 'metadata': {}, 'outputs': [], 'source': 'print(1)', 'idx_': 0, 'lang_': 'python'}]
Notebook files store multiline text as lists of lines, as nbformat’s split_lines produces, in cell sources, stream text, textual MIME data and attachments. As nbformat’s own reader does, reading joins these into plain strings and writing splits them again. In-memory code always sees strings, while files round-trip byte for byte with Jupyter’s. JSON MIME data and error tracebacks are real lists, not split text, and pass through untouched.
disk = dict(nbformat=4, nbformat_minor=5, metadata={}, cells=[
dict(cell_type='code', id='c0', metadata={}, execution_count=1, source=['a=1\n','a'],
outputs=[dict(output_type='execute_result', metadata={}, execution_count=1,
data={'text/plain':['hi\n','there'], 'text/markdown':['single'],
'application/json':{'a':[1,2]}, 'image/png':'iVBOR\nw0KG=='}),
dict(output_type='stream', name='stdout', text=['s1\n','s2\n']),
dict(output_type='error', ename='E', evalue='e', traceback=['t1','t2'])]),
dict(cell_type='markdown', id='m0', metadata={}, source=['# t\n','x'],
attachments={'im.txt':{'text/plain':['a\n','b']}, 'im.png':{'image/png':'aGk='}})])
tmp = Path('tmp.ipynb')
try:
tmp.write_text(dumps(disk))
nb = read_nb(tmp)
res,strm,err = nb.cells[0].outputs
test_eq(res['data']['text/plain'], 'hi\nthere') # textual data joined on read
test_eq(res['data']['text/markdown'], 'single') # one-element lists too
test_eq(res['data']['application/json'], {'a':[1,2]}) # JSON mimes untouched
test_eq(res['data']['image/png'], 'iVBOR\nw0KG==')
test_eq(strm['text'], 's1\ns2\n') # stream text joined
test_eq(err['traceback'], ['t1','t2']) # tracebacks stay lists
test_eq(nb.cells[1].attachments['im.txt']['text/plain'], 'a\nb')
test_eq(nb2dict(nb), disk) # splitting on save restores the disk form exactly
finally: tmp.unlink()diff_cells compares two versions of a cell sequence, such as an open notebook and its file on disk, or a dialog and an edited copy. It yields the edits that turn one into the other. It aligns items by id, and works with any objects that have an .id, not just cells. Each edit is a block covering a contiguous run of deletions, insertions or in-place changes. An applier can handle each block in one step, keeping inserted items in order. An aligned pair counts as changed when key(item) differs between the two sides.
Yield block CellEdits that turn a into b, aligned by item id
Here are five cells, and a copy with one deleted, one edited and one inserted. Edits come in the order of a. old and new are lists, matched item by item for a change. An insert’s idx anchors its block in a: the new items go after a[idx-1], or at the front when idx is 0:
a = [mk_cell(f'x={i}', id=c) for i,c in enumerate('abcde')]
b = copy.deepcopy(a)
del b[1]
b[1].source = 'x=99'
b.insert(3, mk_cell('y=0', id='f'))
edits = list(diff_cells(a, b))
d,c,i = edits
test_eq(d.op, 'delete'); test_eq([o.id for o in d.old], ['b'])
test_eq(c.op, 'change'); test_eq(c.new[0].source, 'x=99')
test_eq(i.op, 'insert'); test_eq(i.idx, 4)
editsBy default key compares whole items, which for cells means the full dict, outputs included. Pass a projection to compare only the content that matters. With key=attrgetter('source'), a difference only in outputs is no edit at all. A projection also handles items that don’t compare by content, such as aidialog messages, whose == compares ids, by mapping each item to a comparable form:
Editing cell dicts directly can produce notebooks that Jupyter and other tools reject, such as a markdown cell with an outputs key, a code cell without execution_count, or duplicate ids. validate_cell and validate_nb are cheap structural checks for those mistakes. They raise a ValueError naming the offending cell and never repair anything, leaving any fix to the caller. They check only the rules whose violation breaks notebooks in practice, not the full nbformat schema.
Raise ValueError for structural problems in notebook nb; returns it unchanged if fine
Raise ValueError for structural problems in notebook cell dict cell; returns it unchanged if fine
A valid notebook passes through unchanged. Each rule fails loudly, naming the cell. The markdown cell with outputs below is a real case, caused by stray keys in a hand-edited file:
vnb = new_nb([mk_cell('1+1'), mk_cell('a note', 'markdown')])
test_eq(validate_nb(vnb), vnb)
vnb.cells[1]['outputs'] = []
with expect_fail(ValueError, 'not allowed in a markdown cell'): validate_nb(vnb)
del vnb.cells[1]['outputs'], vnb.cells[0]['execution_count']
with expect_fail(ValueError, 'requires execution_count'): validate_nb(vnb)dup = new_nb([mk_cell('a', id='x1'), mk_cell('b', id='x1')])
with expect_fail(ValueError, 'duplicate cell id'): validate_nb(dup)
with expect_fail(ValueError, 'unknown cell_type'): validate_cell(dict(cell_type='wat', source=''))
with expect_fail(ValueError, 'str or list of str'): validate_cell(dict(cell_type='raw', source=[1,2]))Fix deterministic structural problems in nb, returning a list of repairs made
Fix deterministic structural problems in cell, returning a list of repairs made
repair_nb applies a deterministic repair for each validation rule:
outputs and execution_count keys.source and metadata values are coerced.It leaves unknown cell_types alone, because no repair can know what was intended. It returns the list of repairs it made: [] for a valid notebook, which it leaves untouched.
bad = dict2nb(dict(cells=[
dict(cell_type='markdown', source='hi', outputs=[], execution_count=1, id='x1'),
dict(cell_type='code', source='1+1', id='x1'),
], metadata={}, nbformat=4, nbformat_minor=5))
with expect_fail(ValueError): validate_nb(bad)
repairs = repair_nb(bad)
validate_nb(bad)
test_eq(len(repairs), 7)
assert all(isinstance(c.metadata, dict) for c in bad.cells)
assert bad.cells[0].id != bad.cells[1].id
test_eq(repair_nb(bad), [])Code outputs need the right structure too. Each entry must be a dict with one of the four standard output_types and that type’s required fields:
stream: name and textdisplay_data: a data dict and metadataexecute_result: a data dict, metadata and execution_counterror: ename, evalue and tracebackRepair fills in a missing metadata or execution_count, and removes entries too malformed to keep, reporting each removal with its reason.
onb = new_nb([mk_cell('1+1')])
onb.cells[0].outputs = [dict(output_type='execute_result', data={'text/plain':['2']}, metadata={}, execution_count=1)]
test_eq(validate_nb(onb), onb)
onb.cells[0].outputs = [dict(output_type='wat')]
with expect_fail(ValueError, 'unknown output_type'): validate_nb(onb)
onb.cells[0].outputs = [dict(output_type='stream', name='stdout')]
with expect_fail(ValueError, 'text'): validate_nb(onb)
onb.cells[0].outputs = [dict(output_type='execute_result', data={})]
with expect_fail(ValueError, 'metadata'): validate_nb(onb)
onb.cells[0].outputs = [dict(output_type='error', ename='E', evalue='boom', traceback='not a list')]
with expect_fail(ValueError, 'traceback'): validate_nb(onb)rnb = new_nb([mk_cell('1+1')])
rnb.cells[0].outputs = [
dict(output_type='execute_result', data={'text/plain':'2'}),
dict(output_type='stream', name='stdout', text='hi'),
dict(output_type='error', ename='E', evalue='boom'),
'junk']
repairs = repair_nb(rnb)
validate_nb(rnb)
test_eq(len(rnb.cells[0].outputs), 2)
test_eq(rnb.cells[0].outputs[0]['metadata'], {})
test_eq(rnb.cells[0].outputs[0]['execution_count'], None)
test_eq(repair_nb(rnb), [])
repairs['cell 58d10327 output 0: set metadata',
'cell 58d10327 output 0: set execution_count',
'cell 58d10327 output 2: removed output (error requires a traceback list of str)',
'cell 58d10327 output 3: removed output (output must be a dict)']
preferred_out selects the best MIME type from an output’s data dict, preferring HTML by default:
data = dict(text_plain=['42'], **{'text/html': ['<b>42</b>'], 'text/plain': ['42']})
wdata = {'image/webp': 'AAAA', 'text/plain': ['im']}
test_eq(preferred_out(wdata, include_imgs=True)[0], 'image/webp')
test_eq(preferred_out(wdata)[0], 'text/plain')
preferred_out(data), preferred_out(data, html1st=False)(('text/html', ['<b>42</b>']), ('text/html', ['<b>42</b>']))
Helper to create an error output dict
Helper to create a display_data output dict
Helper to create an execute_result output dict
Concatenate stream outputs by name (stdout/stderr), preserving execute_result at end
concat_streams merges consecutive stream outputs by name and moves execute_results to the end, as standard Jupyter output rendering does:
[{'output_type': 'execute_result',
'data': {'text/plain': ['42']},
'metadata': {}},
{'output_type': 'stream', 'name': 'stdout', 'text': 'hello '},
{'output_type': 'stream', 'name': 'stdout', 'text': 'world\n'},
{'output_type': 'stream', 'name': 'stderr', 'text': 'warn\n'}]
[{'output_type': 'stream', 'name': 'stdout', 'text': 'hello world\n'},
{'output_type': 'stream', 'name': 'stderr', 'text': 'warn\n'},
{'output_type': 'execute_result',
'data': {'text/plain': ['42']},
'metadata': {}}]
Carriage returns overwrite across chunk boundaries, as on a live terminal. The next chunk overwrites a progress line that ends in \r, but a final \r, with nothing after it, leaves its text visible:
Preferred mime type and content for any Jupyter output dict (stream, error, or data-bearing)
preferred_msg_out extends preferred_out to any output dict. Streams and errors are always plain text, while outputs with data go through MIME preference:
test_eq(preferred_msg_out(mk_stream('stdout', ['a\n','b\n'])), ('text/plain', 'a\nb\n'))
test_eq(preferred_msg_out(mk_result(text_html=['<b>4</b>'], text_plain=['4'])), ('text/html', ['<b>4</b>']))
test_eq(preferred_msg_out(mk_result(text_markdown=['*4*'], text_html=['<b>4</b>']), html1st=False)[0], 'text/markdown')<pre class="!border-0 !rounded-none !my-0 !p-0"><code class="nohighlight">42</code></pre>
Render a full list of outputs, concatenating streams first.
<pre class="!border-0 !rounded-none !my-0 !p-0"><code class="nohighlight">hello world
</code></pre>
<pre class="!border-0 !rounded-none !my-0 !p-0"><code class="nohighlight">warn
</code></pre>
<pre class="!border-0 !rounded-none !my-0 !p-0"><code class="nohighlight">42</code></pre>
Render notebook outputs to concise ANSI-stripped text, using XML-ish tags when multiple outputs are present; tb_maxlen caps over-long error-traceback lines
A single output renders as plain text directly:…
…but multiple outputs are each wrapped in XML-style tags to keep them apart. Rich outputs use markdown by default, or HTML with html1st=True. Outputs with no text form, such as a lone image, render as empty.
<stdout>
hello world
</stdout>
<stderr>
warn
</stderr>
<execute_result>
42
</execute_result>
render_text always strips ANSI escape sequences, because IPython colors its errors. tb_maxlen caps over-long traceback lines, such as the one where a cell magic’s transformed source echoes its whole payload. Capping keeps File and Cell locations whole, along with the final chunk, which holds the exception message. It drops the caret rows that mark positions, since they mean nothing once a line is cut.
tb = ['\x1b[31mFile /a/b.py:1\x1b[0m\n' + 'x'*200 + '\n~~~~^^^' + '~'*200, 'ValueError: ' + 'y'*200]
t = render_text([mk_error(tb, 'ValueError', 'boom')])
test_eq('\x1b' in t, False)
assert 'x'*200 in t
t2 = render_text([mk_error(tb, 'ValueError', 'boom')], tb_maxlen=120)
assert 'x'*200 not in t2 and 'File /a/b.py:1' in t2 and '~~~~^^^' not in t2 and 'y'*200 in t2
print(t2)view_cell can include outputs. incl_out=True appends the rendered outputs in an <out> block, capped at about 512 characters unless trunc_out=False.
vnb = Path('tmp_out.ipynb')
vouts = [dict(output_type='execute_result', metadata={}, data={'text/plain': ['1']}, execution_count=1)]
write_nb(new_nb([mk_cell('a=1\nprint(a)', outputs=vouts)]), vnb)
vcid = read_nb(vnb).cells[0].id
test_eq(str(view_cell(vnb, vcid, incl_out=True)), '1: a=1\n2: print(a)\n<out>\n1\n</out>')
assert '<out>' not in str(view_cell(vnb, vcid))
vnb.unlink()def item2xml(
typ, # Tag name: the cell or message type, e.g. 'code', 'markdown', 'raw', 'prompt'
content:str='', # The item's source text
out:str='', # Rendered output text
id:NoneType=None, # Optional id attribute
meta:NoneType=None, # Cell/message metadata: directives in its `nbdev` dict render as attrs, bare ones as bare attrs
**attrs
):A notebook cell or dialog message as concise XML: content, then an <out> section when out is non-empty
item2xml renders notebook cells and dialog messages as XML for LLMs. cell2xml below and aidialog’s message renderers are built on it. An item’s content sits directly inside its type tag with no wrapper, and any output follows in an <out> section. Passing meta renders the metadata’s nbdev directives as attributes, the same way in every renderer built on item2xml.
'<markdown id="cd"># hi</markdown>'
Falsy attrs are dropped, literal True attrs have no value:
Convert notebook cells to XML format
Convert NbCell to concise XML format
We can view any notebook as concise XML. For instance, our minimal notebook:
<nb><code id="c0">a=1
a</code><markdown id="m0"># t
x</markdown></nb>
We can now open a notebook and access its metadata and cells:
(['solveit_dialog_mode', 'solveit_ver'], 2, 2)
A notebook’s repr is its XML:
<nb path="/Users/jhoward/aai-ws/fastcore/tests/minimal.ipynb"><markdown id="801558df">## A minimal notebook</markdown><code id="e2147a69"># Do some arithmetic
1+1<out>2</out></code></nb>
You can also get a more concise version that doesn’t include outputs or the full path:
<nb path="minimal.ipynb"><markdown id="801558df">## A minimal notebook</markdown><code id="e2147a69"># Do some arithmetic
1+1</code></nb>
Index cells by integer position or by string id, exact or any unique prefix. A missing or ambiguous id raises a KeyError naming the problem:
You can directly set a cell’s source by id or index:
Cells also have the editing operations as methods. They are the cell_* functions without the path and cell id arguments, since they act on the cell you hold. They edit in memory and return a diff. Saving the notebook writes the changes to disk.
@@ -1 +1 @@
-2+2
+3+3
You can also update outputs and metadata directly on a cell:
[{'output_type': 'execute_result', 'data': {'text/plain': ['4']}}]
The add method inserts a new cell at a given position (defaulting to the end):
Add a new cell with source at idx (default: end), or after/before a cell id
(4, '## A minimal notebook')
Cells can also be inserted relative to an existing cell by id:
['# Before first', '## A minimal notebook', '# After first']
Add a new cell with source at idx (default: end), or after/before a cell id
md is a shortcut to add(..., cell_type='markdown')
You can delete by id or index:
Move cells with src_ids after/before a cell id, or to end
Cells can be moved by id, either relative to another cell or to the end:
True
Use save to write to disk:
If no path is passed, the path used in open() will be re-used.
Show cell source with optional line numbers
The view_cell method displays a cell’s source with optional line numbers:
1 │ # Do some arithmetic
2 │ 1+1
<out>
2
</out>
Each query method has a function twin that takes a path in place of the notebook object. The functions return snapshots, not live cells. A CellRow records a cell’s id, type, source and metadata. Its summary line, id:t[directives]:source, where t is c for code, m for markdown or r for raw, shows any nbdev directives, including whether the cell is exported. Rows are for reading and for addressing edits, which go through the cell_* functions or cell_exhash.
One-line preview capped at maxlen, with blank-line runs shown as ¶ and a trailing count of omitted characters
prev_line is the one-line row shared by CellRow and aidialog’s message previews. A row that fits shows its whole text. A truncated row ends with the number of characters omitted from the escaped text:
Built-in mutable sequence.
If no argument is given, the constructor creates a new empty list. The argument must be an iterable if specified.
Snapshot of one cell, shown as id:t[directives]:source (t: c=code m=markdown r=raw); a context row from a find shows - in place of its final :
A long cell’s row ends with the number of omitted characters:
The plain dict form of the held notebook (nb2dict): the representation layer
One snapshot line per cell of the notebook at path
801558df:m:## A minimal notebook
e2147a69:c:# Do some arithmetic\n1+1
find_cells searches cell sources by regex. By default it also returns one neighbouring cell on each side of each match, because in a documented notebook the neighbouring markdown usually explains the match. Results are FoundCells, a plain list with the Found mixin that aidialog’s message finds also use. Index them by cell id, exact or any unique prefix, never by position. Integer indexing raises an error, since fc[0] could be a context row, not a match. In the display, a context row shows - in place of its final :.
Mixin for find-result containers: matched ids, id-only indexing, and match/context kinds
Find results: cells indexed by id (exact or unique prefix), shown as CellRow lines with context rows marked
Find live cells matching all the given criteria, plus context neighbouring cells
The code cell is the match, with its markdown neighbour returned as context and marked with -. Positional indexing fails with an error that explains indexing by id. With context=0, only matches come back:
The path-taking twin returns CellRow snapshots in a FoundCells, with the same indexing contract:
def find_cells(
path, # Notebook file to search
pat:str='', # Regex over cell source
cell_type:str=None, # Optional limit by type ('code', 'markdown', or 'raw')
ids:str='', # Optional limit by cell ids (comma-separated str, or list); exact or unique prefixes
context:int=None, # Cells of context around matches (default 1)
):Snapshot FoundCells for matching cells in the notebook at path
update_cell changes a cell’s attributes rather than lines of its source. Plain keywords such as cell_type= assign, with metadata= replacing the whole dict. mergemeta deep-merges into the metadata. export sets the nbdev export directive. There is no negative export directive. export=True adds the directive and export=False removes it, along with any exports, in whichever form the cell uses, comment or metadata. deep_merge treats a None value as a deletion:
Copy of d updated by u
def update_cell(
path:str, # Notebook file to modify
cell_id:str, # Id of the cell to update (exact, or unique prefix)
mergemeta:dict=None, # `deep_merge` into cell metadata; a `None` value deletes its key
export:bool=None, # Add (True) or remove (False) the nbdev export directive, in whichever form the cell uses
**kwargs
):Update a cell’s attributes, metadata, or export directive
The returned diff covers the cell’s concise XML and its metadata, which between them show every change update_cell can make. Removing the directive from a comment-form cell edits its source. Adding one to a cell without directives writes the comment form too:
Kernels and notebook tools, such as execnb’s CaptureShell.run_all and aidialog’s %nbrun magic, pick code cells from a notebook by cell id. select_cells matches an id or unique prefix. It can also include the cells above or below the match, or all code cells, and can keep only nbdev-exported cells. Whether a selected cell runs follows the eval cascade: the cell’s own eval: directive, else the notebook-level eval directive, else the default_eval argument. A cell named by id always runs, because #| eval: false means “not by default”, not “never”. ignore_eval=True skips the cascade entirely, like run-all in Jupyter:
def select_cells(
nb, # A notebook read with `read_nb`
*msgids:str, # Cell ids, or unique prefixes, to match
above:bool=False, # Include each matched cell and all cells above it?
below:bool=False, # Include each matched cell and all cells below it?
all:bool=False, # Include all code cells (ignores `msgids`)?
exported:bool=False, # Only cells with `#| export` or `#| exports`?
default_eval:bool=True, # Participation default when neither the cell nor the notebook has an `eval` directive
ignore_eval:bool=False, # Skip `eval` filtering entirely: every selected cell runs
):Select code cells from nb by cell id or unique prefix, in the order given; cells named in msgids always run, others follow the eval cascade
Does cell participate in a non-interactive run? Decided by its eval directive, or default when the directive is absent or unrecognized; a cell tagged skip-execution (nbclient’s skip tag) never participates
Participation default for cells without their own eval directive: the notebook-level eval directive in frontmatter mapping fm if given, else default_eval
snb = new_nb([mk_cell('# intro', 'markdown')] + [mk_cell(f'x{i} = {i}', id=f'aa{i}{i}0000') for i in range(4)])
codes = [c for c in snb.cells if c.cell_type=='code']
c1 = codes[1]
test_eq(select_cells(snb, 'aa11'), [c1])
test_eq(select_cells(snb, 'aa110000'), [c1])
test_eq(select_cells(snb, 'aa11', above=True), codes[:2])
test_eq(select_cells(snb, 'aa11', below=True), codes[1:])
test_eq(select_cells(snb, all=True), codes)
test_eq(select_cells(snb, 'aa33', 'aa00'), [codes[3], codes[0]]) # several ids, in the order given
c1.source = '#| export\n'+c1.source
test_eq(select_cells(snb, all=True, exported=True), [c1])
c1.source = '#| eval: false\n'+c1.source
test_eq(select_cells(snb, all=True), [c for c in codes if c is not c1]) # `eval: false` drops out of bulk selection
test_eq(select_cells(snb, 'aa11'), [c1]) # but a cell named by id always runs
codes[2].source = 'nbdev_export'+'()'
test_eq(select_cells(snb, all=True), [codes[0], codes[3]])
test_eq(select_cells(snb, all=True, ignore_eval=True), codes) # Jupyter-style literal run-all
with expect_fail(Exception, 'aa'): select_cells(snb, 'aa')
with expect_fail(Exception, 'msgids'): select_cells(snb)
select_cells(snb, 'aa11')[{'cell_type': 'code',
'source': '#| eval: false\n#| export\nx1 = 1',
'directives_': {},
'id': 'aa110000',
'metadata': {},
'outputs': [],
'execution_count': None,
'idx_': 2,
'lang_': 'python',
'_meta_names_': set(),
'_directives_': {'eval': 'false', 'export': ''}}]
Text that looks like a directive inside a string isn’t a directive. Metadata-form directives count like comment-form ones. A cell tagged skip-execution, nbclient’s own “never run” marker, never runs, matching jupyter execute in both nbdev-test and select_cells.
sc = mk_cell('x = """\n#| eval: false\n"""')
assert sc in select_cells(new_nb([sc]), all=True)
mc = mk_cell('y = 1', metadata=dict(nbdev=dict(eval='false')))
assert mc not in select_cells(new_nb([mc]), all=True)
tc = mk_cell('z = 1', metadata=dict(tags=['skip-execution']))
assert tc not in select_cells(new_nb([tc]), all=True)
assert not does_cell_eval(tc, True) # the tag wins whatever the default: what nbclient skips, nothing runsIn the eval cascade, the notebook-level eval directive can come from any nb_frontmatter source: a frontmatter block, a - key: value line in the title cell, or the notebook’s metadata.nbdev. fm_default_eval resolves that level. It recognizes true, True, false and False as booleans or strings, and ignores anything else. does_cell_eval applies the cell level. With default_eval=False, a run is opt-in: only cells marked #| eval: true run. Template fill uses this to run a dialog’s cells, while batch testing keeps running cells by default:
onb = new_nb([mk_cell('a = 1', id='p1'), mk_cell('#| eval: true\nb = 2', id='m1')])
test_eq(len(select_cells(onb, all=True)), 2) # nothing says otherwise: everything runs
test_eq([c.id for c in select_cells(onb, all=True, default_eval=False)], ['m1']) # opt-in: only marked cells
fnb = new_nb([mk_cell('---\neval: false\n---', 'raw'), *onb.cells])
test_eq([c.id for c in select_cells(fnb, all=True)], ['m1']) # the nb-level directive fills silence
test_eq(fm_default_eval(nb_frontmatter(fnb)), False)
test_eq(fm_default_eval({'eval':'True'}, default_eval=False), True) # four spellings, str or bool
test_eq(fm_default_eval({'eval':'maybe'}), True) # unrecognized: ignoredasync def run_cell(
shell, # An `InteractiveShell`-compatible object: `transform_cell`, `run_cell_async`, `events`
raw_cell:str, # Python/IPython source for one cell: magics, `!` commands, and top-level `await` all work
store_history:bool=False, # Store the cell in the shell's history? (enables native `;` suppression and execution counts)
silent:bool=False, # Suppress displayhook and `post_run_cell` event? (the result value stays unset)
shell_futures:bool=True, # Share `__future__` imports with the shell?
cell_id:NoneType=None, # Optional cell id, passed through to `run_cell_async`
):Run one cell on shell: transform, await run_cell_async on the calling loop, and fire the post events; returns the ExecutionResult
run_cell is the shared execution core. It transforms the source itself, because IPython no longer does that automatically. It then awaits the shell’s own run_cell_async, and fires the post-run events that run_cell_async leaves out. It works with any shell that has transform_cell, run_cell_async and events, such as an InteractiveShell, without fastcore importing IPython. run_cell leaves displays, streams and tracebacks to the shell you pass and the silent flag. A capturing shell records them and a kernel shell publishes them. silent=True suppresses the displayhook and post_run_cell, which leaves the result value unset, while streams still flow. With store_history=True, the shell’s own trailing-; suppression and execution counts work as in a live kernel.
One run returns the value and fires post_run_cell. A failing run records its exception in the result instead of raising. A silent run fires neither the displayhook nor post_run_cell, and records no value. A trailing ; suppresses output as in a notebook:
sh = InteractiveShell()
events = []
sh.events.register('post_run_cell', lambda result: events.append(result))
r = await run_cell(sh, 'y = 6\ny*7', store_history=True)
test_eq((r.result, r.error_in_exec, len(events)), (42, None, 1))
r2 = await run_cell(sh, '1/0', store_history=True)
test_eq((type(r2.error_in_exec), len(events)), (ZeroDivisionError, 2))
rs = await run_cell(sh, '9*9', silent=True, store_history=True)
test_eq((rs.result, len(events)), (None, 2))
r3 = await run_cell(sh, 'y*7;', store_history=True)
test_eq(r3.result, None)
rA fresh InteractiveShell shows the contract: source in, ExecutionResult out, state kept in the shell’s namespace. With store_history=True the shell’s native trailing-; suppression applies, and magics, ! commands, and top-level await all work.
sh = InteractiveShell()
r = await run_cell(sh, 'x = 6\nx*7', store_history=True)
test_eq((sh.user_ns['x'], r.result), (6, 42))
test_eq((await run_cell(sh, 'x;', store_history=True)).result, None)
test_eq((await run_cell(sh, '%who_ls int', store_history=True)).result, ['x'])
(await run_cell(sh, 'import asyncio\nawait asyncio.sleep(0.001)\n"slept"', store_history=True)).resultPrints and displays go through the ordinary user channels, where an enclosing capture sees them. On a plain shell the displayhook also writes the result to stdout, as every non-silent run displays its result. silent=True runs the same code with no display and no result value.
with capture_output() as cap: r = await run_cell(sh, 'print("inner"); 5', store_history=True)
assert cap.stdout.startswith('inner\n') and 'Out' in cap.stdout
test_eq(r.result, 5)
with capture_output() as cap2: rs = await run_cell(sh, 'print("quiet"); 5', silent=True)
test_eq((cap2.stdout, rs.result), ('quiet\n', None))Extensions that display from the execution events work, because the events fire around the cell. matplotlib’s inline backend flushes figures on post_execute, which is all it needs on a bare InteractiveShell with no GUI loop.
Consumers of kernel messages, such as renderers and dialog models, usually want iopub outputs in the form a notebook file stores, without the protocol wrapping. jupywire handles the wire framing itself. msg2out converts one message with the semantics of nbformat.v4.output_from_msg, without depending on nbformat. It copies only the output fields, never the protocol-only transient. It accepts a missing or None metadata, and raises ValueError for a message that isn’t an output. The msg_type can sit at the top level of the message, as in jupyter_client-style dicts, or in the header, as in wire messages. msgs2outs converts a whole drained list, skipping messages that aren’t outputs, such as status and execute_input.
nbformat-style output dicts for the output messages in msgs, skipping other message types
def _hdr_msg(typ, **c): return dict(header=dict(msg_type=typ), content=c) # wire style
def _top_msg(typ, **c): return dict(msg_type=typ, content=c) # jupyter_client style
test_eq(msg2out(_hdr_msg('stream', name='stdout', text='hi\n')), dict(output_type='stream', name='stdout', text='hi\n'))
test_eq(msg2out(_top_msg('error', ename='E', evalue='boom', traceback=['t'])),
dict(output_type='error', ename='E', evalue='boom', traceback=['t']))
test_eq(msg2out(_hdr_msg('execute_result', data={'text/plain':'42'}, metadata={}, execution_count=1, transient={'display_id':'x'})),
dict(output_type='execute_result', metadata={}, data={'text/plain':'42'}, execution_count=1))
test_eq(msg2out(_top_msg('display_data', data={'text/plain':'hi'}, metadata=None, transient={})),
dict(output_type='display_data', metadata={}, data={'text/plain':'hi'}))
with expect_fail(ValueError): msg2out(_hdr_msg('status', execution_state='idle'))
msgs = [_hdr_msg('status', execution_state='busy'), _hdr_msg('execute_input', code='1+1'),
_hdr_msg('stream', name='stdout', text='hi\n'), _top_msg('execute_result', data={'text/plain':'2'}, metadata={})]
msgs2outs(msgs)[{'output_type': 'stream', 'name': 'stdout', 'text': 'hi\n'},
{'output_type': 'execute_result',
'metadata': {},
'data': {'text/plain': '2'},
'execution_count': None}]