Retain only media still referenced by placeholders in current text.
Current input text shown to the user.
Previous input text, used to keep tracking the same placeholder occurrence when duplicate literal tokens are added.
Current cursor offset, used to disambiguate whole-paste edits that create duplicate placeholder tokens.