Skip to content

Base the tokenizer API on source offsets #153569

Description

@pablogsal

The tokenizer currently stores many source positions as pointers into buffers that can move. Reallocating the input means rebasing all these pointers, and missing one is very easy, especially with incremental input and f-strings.

I think tokenizer positions should be offsets into the decoded source instead:

typedef Py_ssize_t TokenizerOffset;

typedef struct {
    TokenizerOffset start;
    TokenizerOffset end;
} TokenizerSpan;

The main ideas would be:

  • One source object owns the decoded text.
  • The cursor only stores its current offset and line boundaries.
  • Tokens, errors and f-string state use offset spans instead of pointers.
  • pegen and _tokenize ask the source for a view or copy of a span.
  • Sequential tokenization keeps the current line in the cursor, while uncommon line lookups can use a small sparse index.
  • Normal parsing can keep one contiguous buffer, while tokenize(readline) can eventually use reclaimable chunks.

For incremental tokenization, new decoded input would be appended only when the cursor needs more data. Existing offsets would remain valid even if the underlying storage moves. Once tokenize has returned copied token and line strings, old input could be discarded when no active token, cursor, f-string frame or error still refers to it.

This is very nice because it separates source storage from tokenizer state, removes pointer rebasing, and means consumers no longer need to access tokenizer internals.

Linked PRs

Activity

  1. Arbaaz123676 commented on Jul 11, 2026

    @Arbaaz123676

    Hi @pablogsal, I'd like to work on this issue. Can you please assign it to me?

  2. pablogsal commented on Jul 11, 2026

    @pablogsal
    MemberAuthor

    Hi @pablogsal, I'd like to work on this issue. Can you please assign it to me?

    Thanks for the help @Arbaaz123676! Unfortunately this is something I am working on right now as a series of commits and is very delicate work. If you want to help reviewing the PRs that would be good!

  3. added 3 commits that reference this issue on Jul 11, 2026
  4. added 3 commits that reference this issue on Jul 18, 2026
  5. added 5 commits that reference this issue on Aug 25, 2026
  6. 220 remaining items

  7. added 2 commits that reference this issue on Sep 29, 2026
  8. added 12 commits that reference this issue on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions