Acknowledgements & Refs
- Vulnerability security advisory: https://github.com/tinymce/tinymce/security/advisories/GHSA-q742-qvgc-gc2f
- TinyMCE patch release notes: https://www.tiny.cloud/docs/tinymce/latest/8.5.1-release-notes/#fixed-stored-xss-vulnerability-through-data-mce-prefixed-src-href-style-attributes
- Affected packages: https://security.snyk.io/vuln?search=cve-2026-47759
Background: What Is TinyMCE?
TinyMCE is a rich text editor that runs inside a web browser. You have probably used it without knowing as it is embedded behind the text boxes on platforms multiple major platforms as it just does its thing behind the scenes.
Be sure to checkout part one of this series to get the initial synopsis of the vulnerability.
Let me preface this by saying this issue was discovered due laziness on my part as I would rather copy and paste that type things out every time. Typing my payload found did nothing as I usually have a specific payload I inject in random parts of an application as I begin exploring it.. Pasting a similar payload led to a rabbit whole that goes beyond CVE-2026-47759.
This is the story of story of CVE-2026-47759, a stored XSS in TinyMCE that lived in versions before 8.5.1. If you had sat there and carefully typed the malicious HTML into the editor character by character, the editor would have quietly stripped everything. But when I pasted a similar payload string from my clipboard, it sailed straight past the one place that was supposed to sanitize it. Big up to implicit trust.
I want to walk you through why and how this happened. How “typing” vs “pasting” gave a payload two completely different fates.
Okay. On to the actual bug.
The bug in one breath
TinyMCE keeps a private set of attributes for its own bookkeeping, all prefixed data-mce-. These include data-mce-href and data-mce-src. The editor uses these to stash the “real” version of a URL while the browser is allowed to mangle the visible one.
The problem is that TinyMCE trusted this channel on the way out without checking who filled it on the way in. It assumed that if an input node has a data-mce-href, then TinyMCE itself must have put it there earlier, from a value it already sanitized. If your raw HTML arrived with an attribute such as data-mce-href already set, the editor carried it along, and on output copied it into the real href without running it through the sanitizer. There goes that implicit trust.
So the attack was simple. Smuggle the payload in through the internal attribute the editor trusts (data-mce-*), and pair it with an innocent-looking public attribute. An example payload would look something like the following:
<a data-mce-href="javascript:alert(1)" href="about:blank">link</a>Two attributes. A boring, harmless href="about:blank" that any sanitizer waves through, and a data-mce-href carrying the real javascript: payload. More how this conversion happens later.
Two doors into the editor, and only one of them parses HTML
Here is the thing that makes this a paste bug specifically, and it comes down to a design fact about how MIME types are interpreted.
When you type into a rich text editor, you are typing into a contenteditable region and the characters are inserted as text rather than interpreted as HTML. If you type an angle bracket, it gives you a literal < character sitting in a text node. It never treats what you typed as markup because isn’t necessarily markup.
Pasting is different because the clipboard may include a text/html representation alongside plain text. If the browser or editor accepts that HTML representation, it parses it into DOM nodes. This all depends on the source you copy from though. If you copy from a plain text source, your clipboard holds text/plain, if you copy from a rich text source, your clipboard holds text/html
A media type (formerly known as a Multipurpose Internet Mail Extensions or MIME type) indicates the nature and format of a document, file, or assortment of bytes. Source: https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/MIME_types
This paste vulnerability occurred because dangerous pasted HTML survived the editor’s filtering and it was transformed into unsafe markup, even though typing the same payload only produced literal text.
On MAC, the following shows the difference in the clipboard when content is copied from a plain text source versus from a rich text source such as Obsidian.
// when you have plain text content copied to clipboard
swift -e 'import AppKit; NSPasteboard.general.types?.forEach { print($0.rawValue) }'
public.utf8-plain-text
NSStringPboardType
// when you have rich text content copied to clipboard
swift -e 'import AppKit; NSPasteboard.general.types?.forEach { print($0.rawValue) }'
public.html
Apple HTML pasteboard type
public.utf8-plain-text
...[snip]...Let me follow the pasted payload through, hop by hop.
Hop 1: the clipboard hands over your HTML
paste/Clipboard.ts:247 registers the listener:
editor.on('paste', (e: EditorEvent<ClipboardEvent>) => {
...
const clipboardContent = getDataTransferItems(e.clipboardData);
...
if (hasContentType(clipboardContent, 'text/html')) {
...
insertClipboardContent(editor, clipboardContent, clipboardContent['text/html'], plainTextMode, true);The editor reads the clipboard, notices it carries text/html, pulls that raw string out, and passes it forward. Nothing has looked at the contents yet. A payload such as data-mce-href="javascript:alert(1)" is still just text in a string at this point.
Hop 2: the pre-sanitizer that does not know your attribute is special
Before the content reaches the main pipeline, paste runs it through its own temporary cleaner in paste/ProcessFilters.ts:13:
const preProcess = (editor: Editor, html: string): string => {
const parser = DomParser({
sanitize: Options.shouldSanitizeXss(editor),
...
}, editor.schema);
// Strip meta elements
parser.addNodeFilter('meta', (nodes) => {
Tools.each(nodes, (node) => {
node.remove();
});
});
...
};Read what this parser is and what it is not. It is a brand new DomParser, created just for paste, with XSS sanitization turned on. It strips <meta> tags. It is not editor.parser, and more importantly it has no idea that data-mce-* attributes mean anything. To this parser, data-mce-href is just some data attribute, and data attributes are ordinary and allowed. The value it carries, javascript:alert(1), is not sitting in a place the XSS sanitizer inspects, because the visible href here is the safe about:blank. The dangerous string is hiding in an attribute the sanitizer has no reason to distrust.
So the payload survives the one component in the whole flow whose function is sanitization. The payloads flies by because the function implicitly trusts that those attributes are safe.
Hop 3: back into the shared parser, the one with the bug
After pre-processing, the content re-enters the normal insertion path. paste/SmartPaste.ts:15 calls editor.insertContent:
const pasteHtml = (editor: Editor, html: string): boolean => {
editor.insertContent(html, {
merge: Options.shouldPasteMergeFormats(editor),
paste: true
});That call threads through the insert command chain and lands in content/InsertContentImpl.ts, where at line 263 it grabs the shared parser and at line 306 it parses:
const parser = editor.parser;
...
const fragment = parser.parse(value, parserArgs);editor.parser is not a fresh one this time. It is the single instance built once at InitContentBody.ts:470. It is the same parser that editor.setContent() uses at SetContentImpl.ts:68. Everything that sets content in the editor funnels through this one object, and this object carries the vulnerable filter.
Hop 4: the filter that trusts the attacker
Here is the heart of it. InitContentBody.ts:143 builds the parser, and in 8.5.0 the first thing it registers is this filter (line 147):
// Convert src and href into data-mce-src, data-mce-href and data-mce-style
parser.addAttributeFilter('src,href,style,tabindex', (nodes, name) => {
const dom = editor.dom;
const internalName = 'data-mce-' + name;
let i = nodes.length;
while (i--) {
const node = nodes[i];
let value: string | null | undefined = node.attr(name);
// Add internal attribute if we need to we don't on a refresh of the document
if (value && !node.attr(internalName)) {
...
node.attr(internalName, editor.convertURL(value, name, node.name));
}
}
});This filter watches src, href, style, and tabindex. On a normal, clean document its job is fine. It takes a real href, runs it through editor.convertURL, and mirrors the result into data-mce-href for TinyMCE’s bookkeeping.
The whole exploit hinges on one word in the guard on line 157:
if (value && !node.attr(internalName)) {“Only create data-mce-href if one does not already exist.” The unspoken assumption is that a data-mce-href can only already exist because TinyMCE put it there on an earlier pass. That assumption is the bug. Your pasted HTML arrived with data-mce-href already set. So node.attr(internalName) is truthy, the condition is false, and this block does nothing. The filter looks at your poisoned internal attribute, decides it must be legitimate prior work, and leaves it completely untouched. It never sanitizes it, because it never believes it needs to.
And notice what is missing in 8.5.0. There is no earlier filter stripping incoming data-mce-* attributes. This src,href,style,tabindex filter is the first one registered in createParser. Nothing ran before it to clean the attacker’s channel.
Hop 5: the payload sits quietly, stored
At this point your malicious data-mce-href="javascript:alert(1)" is a real attribute on a real anchor node in the parsed document. And it does absolutely nothing. It does not render. It does not execute. data-mce-href is not a real link target, it is inert bookkeeping data.
That is exactly why this is stored XSS and not something that fires immediately. The payload survives in the document state, and in a real deployment it survives in whatever database holds the saved editor content. A low-privilege user could plant this in a comment or a profile, save it, and walk away. Nothing looks wrong.
Hop 6: output springs the trap
The detonation happens on the way out from storage. Whenever the content gets serialized back to an HTML string, on getContent, on save, or on re-rendering for the next viewer, a second parser runs the filters in dom/DomSerializerFilters.ts:25:
htmlParser.addAttributeFilter('src,href,style', (nodes, name) => {
const internalName = 'data-mce-' + name;
const urlConverter = settings.url_converter;
const urlConverterScope = settings.url_converter_scope;
let i = nodes.length;
while (i--) {
const node = nodes[i];
let value = node.attr(internalName);
if (value !== undefined) {
// Set external name to internal value and remove internal
node.attr(name, value.length > 0 ? value : null);
node.attr(internalName, null);
} else {
// No internal attribute found then convert the value we have in the DOM
value = node.attr(name) as string;
if (name === 'style') {
value = dom.serializeStyle(dom.parseStyle(value), node.name);
} else if (urlConverter) {
value = urlConverter.call(urlConverterScope, value, name, node.name);
}
node.attr(name, value.length > 0 ? value : null);
}
}
});This is the reverse operation. On output it takes the internal data-mce-href and restores it into the real href. Look closely at the branch on line 35, if (value !== undefined). When a data-mce-href exists, this branch runs, and it copies that value straight into the real href. Then look at what it does not do. It never calls urlConverter. The only sanitization in this whole filter lives in the else branch, the path for values that arrived as a plain href with no internal twin.
So by smuggling the payload in as data-mce-href, the attacker guarantees the if branch runs and the else branch, the one with the sanitization, never does. Your href="about:blank" gets overwritten by javascript:alert(1), unconverted, unchecked. The output HTML now contains a genuinely malicious link.
When that content is rendered and a victim clicks the link, the JavaScript runs in their browser under the application’s origin. Same mechanism applies to src and style, since the filter treats all three identically.
So why does typing miss all of this
Now the payoff, and it is the whole reason you should care about this attack scenario for rich text editors and similar items.
Every hop above started with a raw HTML string going into a parser. Hop 1 pulled a string off the clipboard. Hop 3 fed it to parser.parse. Hop 4 is a parser filter. None of that exists on the typing path, because typing never produces an HTML string for TinyMCE to parse.
If you sit there and type <a data-mce-href="javascript:alert(1)" href="about:blank">link</a> into the editor, the browser’s contenteditable handling inserts those characters as literal text. The angle brackets become the characters < and >, not the start of an anchor tag. There is no anchor node and there isn’t an attribute either. There is nothing for the ingest filter to trust, because you never created an element.
In the 8.5.0 core, the calls to parser.parse come from a small, fixed set of places: the parser’s own definition, InsertContentImpl.ts (insert and paste), SetContentImpl.ts (setContent), the paste pre-sanitizer, and selection-setting. Keystroke and input handling is not on that list. There is no path from a keypress to DomParser.
The architecture is clear enough to state the lesson; if you only ever typed your payloads, in the editor you were using, you would have been testing the one input path that cannot reach the parser. You would poke at the editor, or the application at large and conclude it was safe.
The fix, so you can see the shape of the mistake
For contrast, 8.5.1 (and later versions) adds a filter, registered before the existing one, that nulls out the attacker’s channel at the door for those specific attributes:
parser.addAttributeFilter('data-mce-src,data-mce-href,data-mce-style', (nodes, name) => {
for (let i = 0; i < nodes.length; i++) {
nodes[i].attr(name, null);
}
});That is the entire patch to the ingest side. Any data-mce-src, data-mce-href, or data-mce-style on incoming content gets stripped the moment it is parsed, before the trusting filter ever runs. The output filter in DomSerializerFilters.ts was left untouched, and it did not need touching. After the strip, a data-mce-href can only exist on a node because TinyMCE’s own filter put it there, from a value it just converted. The channel the attacker used is closed at the source instead of at the point where it was trusted.
As I am proof reading this, I realized I actually need to investigate whether there is any way to manipulate TinyMCE’s filter for it to insert malicious content in those attributes; is there any other source basically. Another TODO. If you get there before me, let me know.
The takeaway
Copy and paste your payloads.
Not because pasting is a trick, but because pasting is the input path that runs your string through the parser, and the parser is where attribute-smuggling bugs like this one live. Typing exercises the browser’s text handling. Pasting exercises the application’s HTML handling. If the code handles these differently, that is where we can begin finding cools bugs.
Next time you are testing a rich text editor and the typed payload does nothing, do not close the tab. Put the same payload on your clipboard and paste it in. That single change is the difference between “looks fine to me” and CVE-2026-47759.
Cheers.