A Python binding for simdjson that
parses JSON into native Python objects (dict, list, str, int,
float, bool, None) and serializes them back. It is a drop-in
replacement for the json module's loads, load, dumps and dump, and
it can be over 3 times faster than the standard json.loads and
json.dumps. When you only need part of a document, its lazy parse
function is faster still. It also reads streams of documents (NDJSON, JSON
Lines).
pip install fastsimdjsonWheels are available for Linux, macOS and Windows, for Python 3.10 to 3.14, including free-threaded Python 3.14.
import fastsimdjson
fastsimdjson.loads(b'{"a": [1, 2.5, "x", true, null]}')
# {'a': [1, 2.5, 'x', True, None]}loads(data) accepts bytes, bytearray, memoryview and str, and
returns the same value as json.loads, with the same types and key order.
Integers that do not fit in 64 bits become exact Python ints. NaN,
Infinity and -Infinity are accepted, in any capitalization.
Invalid input raises fastsimdjson.JSONDecodeError, a subclass of
json.JSONDecodeError. When simdjson rejects a document, the same input is
parsed with json.loads. If that succeeds, loads returns its value (this
is how a number that overflows a double becomes inf). If it raises
JSONDecodeError, the exception is re-raised with Python's message and byte
position. Any other exception from json.loads propagates.
When you need only part of a document, parse avoids building the rest.
It accepts the same inputs as loads and returns read-only views:
fastsimdjson.Object (a Mapping) and fastsimdjson.Array (a
Sequence). Values are converted when you access them; nested objects and
arrays are returned as views. A scalar root is returned as a plain value.
doc = fastsimdjson.parse(open("twitter.json", "rb").read())
ids = [(s["id"], s["user"]["screen_name"]) for s in doc["statuses"]]
doc.at_pointer("/statuses/0/user/name") # JSON Pointer (RFC 6901)
doc["search_metadata"].as_dict() # convert a subtree, like loadsObject supports obj[key], get, in, len, iteration over the keys,
keys(), values(), items() (iterators), at_pointer and as_dict().
Array supports arr[i] (negative indexes and slices), len, iteration,
at_pointer and as_list(). Both work with match statements.
- A view keeps its document alive; the document owns its own buffers, so it remains valid while other documents are parsed.
- A key lookup scans the object. With duplicate keys, lookups return the
first value, whereas
as_dict()(likejson.loads) keeps the last. - Indexing an array walks it from the last index reached, so a loop over
arr[i]is linear; iteration is the fastest way to visit an array. - Errors are handled as in
loads. A document that simdjson rejects butjson.loadsaccepts (an overflowing number, an unpaired surrogate) is returned as plain Python objects, asloadswould return it.
fastsimdjson.dumps({"a": [1, 2.5, None]}) # '{"a": [1, 2.5, null]}'
fastsimdjson.dumps(obj, indent=2, sort_keys=True)
fastsimdjson.dumps(obj, separators=(",", ":"), ensure_ascii=False)
with open("out.json", "w", encoding="utf-8") as f:
fastsimdjson.dump(obj, f)dumps(obj, **kw) takes the arguments of json.dumps and returns the same
str, character for character, including the float format (repr), the
escapes and the default separators. ensure_ascii, indent, separators,
sort_keys, allow_nan and default are handled in C. Everything else is
passed to json.dumps itself, which produces the result or raises its usual
exception: a cls argument or other encoder options, skipkeys, a circular
reference, NaN with allow_nan=False, a key or a value that json cannot
serialize. dump(obj, fp, **kw) writes dumps(obj, **kw) to fp.
with open("data.json", "rb") as f:
doc = fastsimdjson.load(f) # like json.load: f.read(), then loads
doc = fastsimdjson.load_file("data.json")
view = fastsimdjson.parse_file("data.json") # lazy, like parseload_file(path) and parse_file(path) accept a str, bytes or
os.PathLike path and raise OSError (e.g. FileNotFoundError) when the
file cannot be read.
for record in fastsimdjson.loads_many(open("log.ndjson", "rb").read()):
...
for view in fastsimdjson.parse_many(data): # lazy views, like parse
...Both return an iterator over the documents of data (bytes, bytearray,
memoryview or str). The format keyword selects how documents are
separated:
format |
input |
|---|---|
"whitespace" (default) |
documents separated by white space, including NDJSON and JSON Lines |
"lines" |
one document per line (NDJSON, JSON Lines) |
"json_seq" |
RFC 7464 JSON text sequences (each document preceded by \x1e) |
"comma" |
documents separated by commas: {...}, {...} |
"array" |
the elements of one array: [{...}, {...}] |
simdjson parses the input in batches (batch_size, 1 MB by default); a
larger document is handled automatically. With "whitespace" and "lines",
documents that simdjson rejects are handled as in loads: a document that
json accepts is returned, otherwise JSONDecodeError reports json's
message and the position in the whole input. A truncated last document is an
error. The views returned by parse_many remain valid after the iterator
moves on.
release() frees the simdjson parser and the string caches kept by the
calling thread. Views returned by parse remain valid.
Python 3.10 or newer, and a C++17 compiler (clang, GCC or MSVC). The simdjson 5.0.2
and simdutf 9.2.1 amalgamations are already in vendor/.
pip:
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -U pip setuptools
python -m pip install -e ".[test]"
pytest testsuv:
uv venv
. .venv/bin/activate
uv pip install -e ".[test]"
pytest testsEither one builds the extension and makes import fastsimdjson work in the
virtualenv. uv uses the setuptools build requirement from pyproject.toml,
so it does not need a separate setuptools install for this path.
To compile the extension in the tree instead, install setuptools and pytest into the same virtualenv, then:
python -m pip install setuptools pytest
python setup.py build_ext --inplace
PYTHONPATH=src python -m pytest testsuv pip install setuptools pytest
python setup.py build_ext --inplace
PYTHONPATH=src python -m pytest testsRecent setuptools copies the .so next to src/fastsimdjson.cpp, which is
why PYTHONPATH=src is required for the in-place build.
tests/test_loads.py compares loads with json.loads on types and key
order; tests/test_lazy.py checks the views returned by parse the same way,
tests/test_dumps.py compares dumps with json.dumps (output and
exceptions), and tests/test_stream.py and tests/test_files.py cover
streams and files. It covers scalars, integers past 64 bits, UTF-8 strings at every
length from 0 to 199, the key cache, random documents, rejected input, deep
nesting, padding at a page boundary, a saturated array count, reference
counts, and release of a parser that has grown past 64 MB. The corpus test
is skipped until simdjson-data is checked out beside the project:
git clone --depth 1 https://github.com/simdjson/simdjson-data.git
pytest tests
# or: JSONDIR=/path/to/jsonexamples pytest testsThe suite builds an ~80 MB document and a list of 16,777,221 integers, so give it some RAM.
To time loads against json.loads and orjson on those files
(bench_lazy.py times parse against pysimdjson and cysimdjson,
bench_dumps.py times dumps, and bench_many.py times loads_many):
python -m pip install -e ".[bench]"
python bench.pyuv pip install -e ".[bench]"
python bench.pyIntel Xeon Gold 6548N (Emerald Rapids), one core, Python 3.14.6, fastsimdjson 0.2.0, the 22 files of simdjson-data. The scripts and the full results are in the blog repository.
Each parser produces the whole document as Python objects. Speed is the geometric mean over the 22 files (higher is better).
| parser | GB/s | vs json.loads |
|---|---|---|
| json (standard library) | 0.22 | 1.00× |
| simplejson 4.1.2 | 0.23 | 1.05× |
| python-rapidjson 1.25 | 0.24 | 1.10× |
| ujson 6.0.0 | 0.36 | 1.65× |
| cysimdjson 26.27 | 0.43 | 1.94× |
| pysimdjson 7.0.2 | 0.44 | 1.98× |
| msgspec 0.22.0 | 0.53 | 2.41× |
| orjson 3.12.0 | 0.60 | 2.73× |
fastsimdjson loads |
0.77 | 3.49× |
fastsimdjson is the fastest on 21 of the 22 files; orjson is slightly faster
on numbers.json, an array of floating-point numbers. Part of the gain comes
from pausing the garbage collector while the objects are built: if the
collector is disabled for every parser, fastsimdjson's lead over orjson drops
from 1.28× to 1.18×. yyjson 4.0.6 is left out: it returns wrong strings for
non-ASCII text.
Parsing is no longer the bottleneck. simdjson alone parses these files at
3.0 GB/s. It accounts for about a third of the time of loads; the rest goes
into creating Python objects. Freeing those objects later costs about a sixth
of the total. Even if parsing took no time at all, loads would be less than
1.5 times faster.
If you only need a few values, parse creates only those. Extracting the id
and the screen name of the 100 statuses of twitter.json:
| method | µs |
|---|---|
json.loads |
3879 |
| orjson | 1008 |
fastsimdjson loads |
860 |
msgspec (typed Struct) |
336 |
| cysimdjson (lazy) | 235 |
| pysimdjson (lazy) | 183 |
fastsimdjson parse |
155 |
Here parse is 25 times faster than json.loads and 5.5 times faster than
loads. Most of its time is the simdjson parse itself: reading the 200
values takes less than 20 µs. Compared with pysimdjson on other tasks
(µs, lower is better):
| file | task | fastsimdjson parse |
fastsimdjson loads |
pysimdjson |
|---|---|---|---|---|
| open | 140 | 676 | 156 | |
| citm_catalog | open | 369 | 1612 | 472 |
| citm_catalog | extract | 394 | 2142 | 504 |
| gsoc-2018 | open | 615 | 2206 | 813 |
| visit all | 1985 | 2183 | 3033 | |
| canada | visit all | 17528 | 18998 | 19750 |
| twitter_api_response | open | 3.6 | 13.6 | 3.4 |
"open" parses the document and looks at its root; "extract" collects the
start time of every performance; "visit all" walks every value through the
views (with loads: through the dict). When you visit everything, parse
is about as fast as loads. On very small documents (15 KB), pysimdjson's
parse is marginally faster.
Same machine, fastsimdjson 0.3.0, the objects of the 22 files. With the
default arguments, dumps returns exactly what json.dumps returns and is
3.4 times faster (geometric mean; from 2.5 times on text-heavy files to 8
times on files full of numbers). Microseconds:
| file | json.dumps |
fastsimdjson dumps |
orjson | msgspec |
|---|---|---|---|---|
| 1482 | 529 | 199 | 347 | |
| citm_catalog | 2702 | 1071 | 427 | 494 |
| github_events | 158 | 43 | 19 | 30 |
| canada | 38622 | 4843 | 2923 | 3653 |
| numbers | 2515 | 349 | 198 | 332 |
orjson and msgspec are faster still, but they produce something else: they
return bytes, without spaces after separators and without escaping
non-ASCII characters. Compared with separators=(",", ":") and
ensure_ascii=False, the closest dumps settings, orjson is 2.9 times
faster and msgspec 1.8 times faster (geometric means).
20 MB of NDJSON (5268 objects and arrays, one per line), made from the same files:
| method | ms | GB/s |
|---|---|---|
json.loads on each line |
167 | 0.12 |
| orjson on each line | 81 | 0.24 |
fastsimdjson loads on each line |
61 | 0.32 |
fastsimdjson loads_many |
52 | 0.38 |
fastsimdjson parse_many (views only) |
20 | 0.96 |
dumpsreturns astr, likejson.dumps; there is no option to returnbytes.- In streams, a document that is a bare number (
3.14alone on its line) is slow to parse: simdjson copies the rest of the batch for each one. Streams of objects and arrays are not affected. - With the
"json_seq","comma"and"array"stream formats, documents that simdjson rejects raiseJSONDecodeErrorwith simdjson's message; there is no fallback tojson. - The simdjson parser and the key and string caches are thread-local.
release()frees the parser and the cached strings retained by the calling thread. A parser that grows past 64 MB is freed on its own at the end of that call; its caches stay. A thread that exits withoutrelease()leaves its cached strings behind. The module is marked free-threading compatible (Py_MOD_GIL_NOT_USEDon Python 3.13 and newer), so importing it on a free-threaded build does not re-enable the GIL. Subinterpreters are not supported. On a free-threaded build,bytearrayandmemoryviewinputs are copied before parsing. - A document simdjson rejects is reparsed with
json.loads.loadsreturns that value whenjson.loadsaccepts it, which is how overflow to infinity is handled. An exception is raised only whenjson.loadsalso fails.JSONDecodeErroris re-raised asfastsimdjson.JSONDecodeErrorwith the same message, document, and position. Any other exception propagates.