core-jmp core-jmpdeath of core jump

Static Devirtualization of Tencent VM: ACE Kernel Drivers, CET, SEH, and Guided Symbolic Execution

Aftermath Labs statically recover 815 of 865 Tencent-VM virtualized functions across four ACE kernel drivers (94.2 percent). Architecture, CET RDSSPQ/INCSSPQ, phantom SEH unwind, guided symbolic execution, VJCC DAGs and a closed MBA identity list.

oxfemale October 2, 2026 41 min read 91 reads
Export PDF
Static Devirtualization of Tencent VM: ACE Kernel Drivers, CET, SEH, and Guided Symbolic Execution
Original text: “Static Devirtualization of Tencent VM” — author not clearly listed (site: Aftermath Labs), Aftermath Labs (31 July 2026). Code, tables, figures and original looping videos below are reproduced verbatim with attribution captions. A mirror of the same post is also served from back.engineering.

Executive Summary

Tencent’s kernel anti-cheat stack, shipped as a family of ACE-*.sys drivers, protects selected functions by rewriting them into a custom virtual machine. The original entry point becomes a trampoline into a .tvm section. Guest registers and flags are saved into a VM context, a shared dispatcher walks encrypted bytecode, and anything the VM does not model is executed as a boxed native instruction. On paper this looks like a serious barrier: the analyst no longer sees ADD and Jcc, they see an interpreter.

Aftermath Labs’ 31 July 2026 note argues the opposite. Classic virtual-machine obfuscators of this shape are weak against a lifting and recompilation framework that is allowed to know a handful of VM-specific facts. Guided symbolic execution inlines the dispatcher, promotes bytecode fetches to constants, folds a closed list of mixed Boolean-arithmetic identities, rewrites virtualized branches back into native Jcc, and lowers the result with a recovered stack frame. Across four ACE kernel drivers they report 815 of 865 virtualized functions recovered to native code — 94.2 percent — with ACE-GAME.sys at 99.3 percent. The output is clean enough that a third party (Daax of secret club / revers.engineering) was given the files to confirm coverage independently. This draft walks the architecture, the CET and SEH compatibility tricks, the IR rewrite rules published in the source, and the hunting implications for anyone who has to live with VM-protected kernel modules.

Kitchen table: Imagine a restaurant that rewrites every recipe into a private cooking language. The kitchen still produces the same dish, but a visiting chef cannot read the cards. A virtual-machine obfuscator does that to a function: the CPU still computes the original answers, through an interpreter that only the protector understands. Aftermath Labs’ claim is that this particular private language is regular enough that a compiler can translate it back.
For operators: The technical thesis is compiler-shaped, not emulator-shaped. Lift AMD64 to SSA, inline the dispatcher along CALL+int3 edges, promote loads whose address is VIP/.tvm0 bytecode, run inst-combine to a fixed point on sixteen MBA identities, match virtual JCC DAGs, then lower with a stack-frame high-water mark taken at boxed/CALL sinks. RSP is concretized as a sentinel, the same way this group already did for Themida, VMProtect, vxlang and the Denuvo anti-cheat driver. The public artefact is the write-up plus IR patterns, not a drop-in ACE unpacker.
Cinematic illustration of bytecode collapsing into native machine code
Featured illustration for this draft: a virtual-machine bytecode volume being pulled into native traces. Additional explainer produced for this draft.

How to read this piece

Two audiences share the same article. Green boxes keep the mechanics in ordinary language. Blue boxes keep the operator-grade facts: offsets, unwind codes, instruction encodings, coverage numbers, and the exact IR the source published. Original screenshots, looping videos, the coverage table and the three code listings appear in the same reading order as Aftermath Labs placed them. Extra diagrams sit next to those originals; they do not replace them.

  • If you do not reverse binaries for a living: read the green boxes, the architecture animation, the CET receipt-tape metaphor, the coverage table, and the defensive checklist. You will leave knowing what a VM protector is, why Tencent had to talk to Intel CET and Windows SEH, and why “it is virtualized” is no longer a conversation-ending claim.
  • If you lift obfuscators: the blue boxes plus the verbatim VJCC patterns, VJCC template and MBA identity list are the payload. Pair them with the source’s boxed-instruction sink trick and the VIP-in-context heuristic.
  • If you hunt kernel anti-cheat or VM-packed malware: jump to the coverage statistics, the trampoline census (E9 + int3), .tvm / .tvm0 sections, RDSSPQ/INCSSPQ in kernel code, and the detection notes. The same shapes show up in Themida and VMProtect families.

What a virtual-machine obfuscator actually is

A virtual-machine obfuscator does not put the program inside Hyper-V or VMware. It replaces selected native functions with an interpreter. The original x86-64 is encoded as bytecode for a private ISA. A dispatcher fetches the next virtual instruction, a handler implements it against a virtual register file, and control returns to the dispatcher. From the outside you see a small cluster of native handlers and a large ocean of data. From the inside the program still adds, compares and branches.

That design has been the workhorse of commercial protectors for two decades. VMProtect, Code Virtualizer / Themida, and a long tail of in-house VMs all share the same silhouette: save GPRs, enter a dispatch loop, hide the real control flow in bytecode. Tencent VM sits in that family. Aftermath Labs say they have believed for a long time that this classic style is not strong against an attacker who already owns a flexible lifter and recompiler. The Tencent work is the latest public demonstration of that belief, after their Themida write-up earlier in 2026.

Kitchen-table metaphor of a ciphered cookbook translated into a readable recipe
The kitchen-table picture: a private recipe language on the left, the same dish written plainly on the right. A VM obfuscator is the left-hand book. Static devirtualization is the translation. Additional explainer produced for this draft.
Kitchen table: The VM is a waiter who takes your order, walks into a kitchen you cannot see, and brings back the correct plate. You never watch the cook. Devirtualization is hiring a translator who has learned the kitchen’s slang, then writing the original recipe back onto a card you can read. You still do not steal the restaurant. You recover the card so you can reason about what the cook is doing.
For operators: Classic VM protectors buy time by forcing the analyst onto a path explosion: every bytecode stream is a new program, handlers are mixed with junk, and the real CFG is an indirect jump through a virtual instruction pointer (VIP). The counter is to stop treating the dispatcher as a black box. Once VIP is a tracked SSA value and bytecode loads are constants, the interpreter collapses into the guest program. That is the whole game. Everything else in the Aftermath Labs note — CET, SEH, MBA, JCC DAGs, high-water marks — is scaffolding so the collapse produces a function you can actually recompile.

Scope, ACE, and the interoperability frame

The target is Tencent’s ACE kernel anti-cheat, observed as ACE-GAME.sys, ACE-BASE.sys, ACE-BOOT.sys and ACE-CORE.sys. Aftermath Labs frame the work as reverse engineering for interoperability with ACE-protected software on Linux and Proton, under the exceptions that recognize reverse engineering for interoperability. They state that sharing of recovered artefacts was limited to trusted third parties for independent technical validation, and that the original protected software remains under copyright and the DMCA.

That frame matters for how this draft treats the material. The source publishes architecture, unwind tricks, IR patterns and coverage numbers. It does not publish a ready-to-run ACE unpacker, a cheat, or a bypass of a live anti-cheat service. This draft stays inside that envelope. Public IR and MBA identities are reproduced verbatim because they are already on the page. A working driver rewriter is not.

This work constitutes reverse engineering, decompilation, and de-virtualization performed solely for the purpose of achieving interoperability with ACE-protected software on Linux and Proton environments.

Aftermath Labs, legal disclaimer on the 31 July 2026 post
Kitchen table: Anti-cheat is the bouncer at a club. The bouncer’s instruction manual is written in a private shorthand so that people who want to sneak in cannot photocopy it. Researchers who want Linux and Proton to run the same title sometimes have to read that shorthand to make the bouncer and the guest OS shake hands. That is a different job from forging a fake stamp.
For operators: CWE mapping for the protector itself is not a vulnerability ID; this is software protection, not a memory-corruption bug. The relevant ATT&CK technique for the protector’s effect on analysis is T1027.002 (Software Packing) sitting under Defense Evasion, with T1027 more broadly covering obfuscated files or information. Kernel ACE also collides with T1562 (Impair Defenses) when it is viewed from the game-process side, and with T1068 only if an analyst is hunting a buggy ACE driver as an LPE surface — which this write-up does not claim. Keep those buckets separate: packing versus privilege.

Prerequisite material the source points at

Aftermath Labs list three pointers before the introduction. They are not reproduced here; they are the reading list the authors assume.

  • k0mkc, 29 July 2026 — a public Tencent VM note posted two days before this article. The Aftermath Labs piece reads as a more complete static-devirtualization treatment sitting on top of a conversation that was already moving.
  • arXiv:2603.18355 — listed as prerequisite theory. Treat it as the authors’ pointer into the current lifting / MBA literature rather than as a substitute for the Tencent-specific heuristics below.
  • arXiv:1909.01752 — a 2019 reference in the same list. Mixed Boolean-arithmetic identities of the kind Tencent uses have a long academic trail (MBA-Blast at USENIX Security 2021 is the paper most operators already know). The Tencent list is interesting because it is small and closed.

A later public toolchain, jz0/tvm-devirt, credits k0mkc and Back Engineering Labs and implements a static TVM devirtualizer in Rust. That repository is a third-party artefact. This draft does not treat it as Aftermath Labs’ official release, and it does not re-host recovered ACE drivers.

Introduction

Aftermath Labs write that interest in Tencent VM obfuscation has been climbing for months, that they have had complete static devirtualization of it for some time, and that others have reached similar deobfuscation results. They chose to publish once that fact was already in circulation. The article’s job is to explain the VM in detail: how it is built, how it stays compatible with Control-flow Enforcement Technology and Structured Exception Handling, and why it loses to guided symbolic evaluation.

A reader of their earlier Themida article had asked for coverage statistics and versioning. They do not have version information for Tencent VM. They do publish a per-driver census of virtualized versus recovered functions, and they state that Daax from secret club was given the files so the cleanliness of the output can be confirmed independently.

Kitchen table: When several groups can already read the private recipe, publishing the grammar is less a leak and more a map. The map is useful to people who will meet the same VM in malware, in DRM, and in anti-cheat, because the grammar is shared even when the vendor is not.
For operators: The strategic claim to keep in a lab notebook: in-house VM protectors keep failing for the same structural reason. They encode a subset of AMD64, they leave a dispatcher that is a single function, they stash VIP in a context object, and they implement flags as ordinary data. A lifter that is allowed to know those four facts does not need a miracle SMT query. It needs constant promotion, a small MBA ruleset, and the discipline not to unroll virtual loops.

Virtual machine architecture

A virtualized function does not begin with its original prolog. The entry point jumps into the .tvm section. Stack space is allocated for a VM context. Every general-purpose register and EFLAGS is saved into that context — first onto the stack, then moved into the context area. Aftermath Labs note the resemblance to VMProtect, which pushes all GPRs and then has the first handlers pop them into the VM’s own storage. Execution then threads in and out of a shared dispatch loop. The loop is left in order to perform calls and boxed instructions.

Labeled diagram of Tencent VM entry, context, dispatcher, handlers and boxed operations
Guided map of the architecture described in the source: trampoline into .tvm, GPR/EFLAGS save, shared dispatcher, handlers, boxed native ops, and VMEXIT. Additional explainer produced for this draft.
Tencent VM architecture — original looping diagram of entry, context save and the shared dispatcher. Original looping diagram. Source: original article.
Kitchen table: Picture a coat-check. Before you enter the private hall (the VM), you hand over every coat, bag and hat (the CPU registers). The hall has its own coat-check tickets (the VM context). When the hall needs to step back onto the public street to make a phone call (a native CALL or a boxed instruction), it puts your coats back on, does the thing in public, then checks the coats again. That is why a screenshot of a virtualized function looks like a tiny stub: the real work is in the hall.
For operators: Census trick published in the coverage section and already implied here: a virtualized function is identified by an E9 rel32 trampoline padded with int3. After a successful recovery the trampoline is retargeted from the .tvm0 interpreter to recompiled .devirt code. Functions the tool could not recover still point into .tvm0. R11 is later described as holding the bytecode address the interpreter consumes. VIP itself lives in the VM context; the published heuristic finds it by looking for the last store of a pointer into .tvm0.

Boxed instructions

Tencent VM models only a subset of AMD64. Anything outside that subset is a boxed instruction. The VM restores native context, executes the original opcode on the real CPU, then captures register state back into the VM context and resumes dispatch. At the moment the boxed instruction runs, machine state is indistinguishable from what the unvirtualized function would have produced, so the instruction does not need to be modelled at all.

That is a practical necessity, not a flourish. CPUID cannot be faked inside a software VM if you want the real leaf values. RDMSR, WRMSR and other architectural instructions are in the same bucket. More usefully for a lifter, every boxed instruction is a point where guest state is forced to materialize. Aftermath Labs use exactly that property to recover the original stack-frame size.

Boxed instruction path — original looping diagram of context restore, native execute, recapture and resume. Original looping diagram. Source: original article.
Kitchen table: Some kitchen tools cannot be described in the private language. A blowtorch is a blowtorch. When the recipe needs one, the cook steps back into the ordinary kitchen, uses the real tool, and returns. Those visits are gold for a translator: they are the moments when the private language has to admit how the public kitchen is actually arranged, including how much counter space (stack) the dish needs.
For operators: Sink-point list used later in lowering: native CALLs and boxed instructions. At those sites RSP, treated as a concretized sentinel, reveals the original frame’s high-water mark because AMD64 PE functions almost never do dynamic alloca. If you are building a sister lifter, log every VMEXIT, dump guest RSP relative to the entry RSP, and take the maximum. That number is the original frame. Boxed ops also receive chained unwind info, because they run outside the one big .pdata blanket that covers the VM proper.

CET compatibility

Entering the dispatcher with CALL is a problem on hardware that implements Control-flow Enforcement Technology. A CET CALL pushes the return address onto a shadow stack as well as onto the data stack. A RET pops both and compares them. The VM’s calls into the dispatcher never return. Left alone, the shadow stack would grow for the lifetime of the virtualized function and eventually fault.

Tencent VM handles this at runtime rather than at protection time. It executes RDSSPQ to read the current shadow-stack pointer. RDSSPQ lives in the hint-NOP encoding space: on a processor or in a process without shadow stacks it retires as a NOP and leaves the destination register untouched. Zeroing the register beforehand and testing it afterward is therefore a feature check that is correct on old and new hardware and costs nothing on either.

Disassembly of rdsspq r15; cmp r15, 1; jz CET feature check
RDSSPQ used as a live CET / shadow-stack feature check (destination r15, compared against 1). Source: original article.

When the check says CET is active, the VM issues INCSSPQ with a register holding 2. That advances the shadow-stack pointer by two entries and discards the two return addresses that no matching RET will ever consume.

Disassembly of mov r15, 2; incsspq r15
INCSSPQ 2: discard two unmatched shadow-stack entries after CALL into a dispatcher that never returns. Source: original article.
Labeled diagram of unmatched CALLs on the CET shadow stack and INCSSPQ 2
Data stack versus shadow stack. Unmatched CALLs into the dispatcher would grow SSP forever; INCSSPQ 2 throws those two shadow entries away. Additional explainer produced for this draft.
Two receipt tapes as a metaphor for the CET shadow stack
Kitchen picture of CET: a second, hidden receipt tape that must stay in lockstep with the ordinary one. A CALL that never RETs is a receipt the shop never tears off — unless someone advances the tape by hand. Additional explainer produced for this draft.
Kitchen table: CET is a second copy of every “I will come back here” sticky note, kept in a locked drawer the program is not supposed to touch. If you keep walking into a room and sticking a note on the door without ever coming back to peel it off, the drawer fills up and the building’s alarm goes off. Tencent’s VM asks “is the locked drawer even installed?” with RDSSPQ, and if it is, it peels two leftover notes with INCSSPQ 2.
For operators: RDSSPQ / INCSSPQ are the kernel-mode hunting tell. A code section that is already an interpreter, already doing E9 trampolines into .tvm, and also emitting a hint-NOP-space RDSSPQ followed by a conditional INCSSPQ 2 is waving a flag. On hardware without CET the same bytes are cheap NOPs, so the check is always present. Do not signature the immediate 2 alone; signature the pairing with a prior RDSSPQ and a dispatcher that is reached by CALL and never RET. Intel’s SDM documents both instructions. kCET in the Windows kernel is the sibling story for kernel CET; this VM is doing user-of-CET hygiene inside a kernel driver’s own call convention.

SEH compatibility

The VM is covered by a single large .pdata entry whose unwind information spans the entire VM range: entry, dispatcher and handlers. That unwind info carries a language-specific exception handler for the VM. A never-executed dummy function carries what Aftermath Labs call phantom unwind info. Its only purpose is to describe unwind operations.

VM entry reserves a 0x68-byte local area. A 0x48-byte block of that is used exclusively during unwinding and holds all nine non-volatile registers. The phantom unwind info describes those saves. The address of the dummy function is kept in [RBP+0] for the entire lifetime of VM execution. The entry’s unwind info carries UWOP_SET_FPREG with RBP as the frame register, so unwinding is performed relative to RBP. RBP is saved at entry and is never used as scratch during VM execution, so the unwinder can recover a frame at any point.

When an exception is raised inside the VM, the VM’s exception handler copies every guest non-volatile register into the save area reserved at entry, overwrites the unwind target slot located eight bytes below the guest RSP with the guest RIP, and returns ExceptionContinueSearch.

Disassembly of TVMExceptionHandler staging non-volatile registers and guest RIP
TVMExceptionHandler: guest non-volatiles are staged with stosq from the VM context, guest RIP is planted, then the handler returns 1 (ExceptionContinueSearch). Source: original article.

The unwinder, following [RBP+0], picks up the dummy function as the next entry. Its .pdata is looked up. Through that phantom unwind info every guest non-volatile the handler just staged is written back into the native register of the CONTEXT record it belongs to. The unwinder then reads the unwind-target slot as RIP and advances RSP by eight, so CONTEXT->RIP holds the guest RIP and CONTEXT->RSP lands exactly on the guest RSP. The original function’s virtual unwind continues as guest state. Boxed instructions get chained unwind info as needed because they run outside the VM.

Four-step diagram of VM exception handling and phantom unwind info
The four hops: reserve and freeze RBP, stage guest non-volatiles and RIP, ExceptionContinueSearch, phantom unwind into CONTEXT. Additional explainer produced for this draft.
Kitchen table: SEH is the building’s fire-escape map. If a kitchen fire starts in the private hall, Windows still has to know which doors to open and which coats to hand back so the rest of the building can unwind the stack as if the private hall had been an ordinary room. Tencent forged a map of a room that is never actually walked through (the dummy function) and taped the real coat locations onto it at fire time. The fire department follows the dummy map and ends up in the right place.
For operators: The TVMExceptionHandler screenshot is doing exactly the write-up’s prose. Context pointer comes from [r8+0A0h] / [rdi+0A0h]; non-volatiles are issued with stosq from fixed VM-context offsets (0x68, 0x70, 0x78, 0x80, 0x30, 0x10, 0x40, 0x38, 0x28 with a sub 8 on one of them); then it writes the staged RIP through [rax] after rcx := [rcx] + 0, and returns 1. The +8 slot below guest RSP is the extra “return address” the unwinder will pop, which is why planting guest RIP there yields CONTEXT->RSP == guest RSP after that pop. This is also why crash dumps of a Tencent-virtualized kernel path can look sane: RtlVirtualUnwind is being fed a lie that is carefully consistent.

Guided symbolic execution

Guided symbolic execution, as used here, means lifting native AMD64 to a higher-level SSA intermediate representation that can be optimized and rewritten. The objective is to symbolically evaluate the entire virtualized function so that one IR function holds all of the routine’s semantics. The lifting loop needs guidance when it hits indirect control flow. Classical VM obfuscation uses bytecode to choose which handler runs next. In this VM, register R11 holds the address of the bytecode the interpreter consumes.

Guided symbolic execution — original looping diagram of lifting and inlining the dispatcher. Original looping diagram. Source: original article.

The guidance is obfuscation-specific. For Tencent VM the authors want to symbolically inline the call to the dispatcher loop. A blunt heuristic works: follow calls that have an int3 placed directly after them. In the symbolic-evaluation path, every one of those calls is a call into the dispatcher. The alternative is to declare the dispatcher function as a valid call target to follow or inline.

Indirect control flow inside the VM is solved by promoting bytecode to a constant inside the SSA IR, so later passes can fold the decryption arithmetic away. That promotion must be scoped. Original semantic loads in the guest program must not be frozen into constants. If they are, the recovered function will hallucinate a single memory image.

Constant promotion — original looping diagram of turning VM bytecode loads into SSA constants. Original looping diagram. Source: original article.

Virtualized loops must not be unrolled by relifting the same VIP forever. The lifter tracks VIP and, if the next VIP / handler has already been lifted, inserts a back-edge. VIP tracking is obfuscation-specific. For Tencent VM, VIP lives in the VM context. The offset is resolved dynamically by finding the last stored value into the context that is a pointer into .tvm0 (bytecode lives there). Aftermath Labs say the heuristic works well and reveals VIP automatically.

Six-stage guided symbolic execution pipeline from lift to lowering
The pipeline as an operator checklist: lift, guide, promote, fold, rewrite VJCCs, lower. Additional explainer produced for this draft.
Kitchen table: Symbolic execution is doing algebra with a program instead of with x and y. “Guided” means the algebra teacher is allowed to whisper, “that indirect jump is just the waiter reading the next line of the private recipe.” Once the next line is a real number (a constant bytecode address), the waiter disappears and you are looking at the recipe.
For operators: Two failure modes to instrument when you clone this. First, over-promotion: a guest load from a stack slot that happens to look like a .tvm0 pointer must stay a load. Scope the promotion to addresses computed from the VIP slot, or to reads whose base is the bytecode mapping. Second, under-guidance: if you miss a dispatcher CALL that is not followed by int3, the lifter will treat the dispatcher as an external function and the IR will be an empty shell. Aftermath Labs offer the int3 heuristic as sufficient on this VM; a second oracle (a known dispatcher RVA, or a handler-table xref) is cheap insurance.

Virtualized conditional control flow

Native Jcc is a single instruction that reads EFLAGS and picks a target. Tencent VM expands the flag comparison into multiple VM handlers. When symbolic evaluation stops on indirect control flow whose destination is still symbolic, two things are possible: you are sitting on a virtualized Jcc, or your optimizations are incomplete.

Virtualized Jcc logic is turned back into a native Jcc by matching pre-defined IR SSA DAGs. If the current indirect control flow matches a DAG, a rewrite fires. Branch targets come out of the same DAG, because the pattern also says where the destinations live in the IR.

The source publishes the pattern table first: one family per flag, four patterns each, covering whether the handler subtracted zero or subtracted the mask, and whether the subsequent test is equality or inequality on ZF of that subtraction. Then it publishes the generic template that those patterns instantiate.

; ── CF ── mask 0x1, idx 0
pattern vjcc.cf.ae.zero { body(0x1,   0x0)   ; %r = R{e}  %z } => { %r = R{ae} %cf }
pattern vjcc.cf.b .zero { body(0x1,   0x0)   ; %r = R{ne} %z } => { %r = R{b}  %cf }
pattern vjcc.cf.b .mask { body(0x1,   0x1)   ; %r = R{e}  %z } => { %r = R{b}  %cf }
pattern vjcc.cf.ae.mask { body(0x1,   0x1)   ; %r = R{ne} %z } => { %r = R{ae} %cf }

; ── PF ── mask 0x4, idx 1
pattern vjcc.pf.np.zero { body(0x4,   0x0)   ; %r = R{e}  %z } => { %r = R{np} %pf }
pattern vjcc.pf.p .zero { body(0x4,   0x0)   ; %r = R{ne} %z } => { %r = R{p}  %pf }
pattern vjcc.pf.p .mask { body(0x4,   0x4)   ; %r = R{e}  %z } => { %r = R{p}  %pf }
pattern vjcc.pf.np.mask { body(0x4,   0x4)   ; %r = R{ne} %z } => { %r = R{np} %pf }

; ── ZF ── mask 0x40, idx 3
pattern vjcc.zf.ne.zero { body(0x40,  0x0)   ; %r = R{e}  %z } => { %r = R{ne} %zf }
pattern vjcc.zf.e .zero { body(0x40,  0x0)   ; %r = R{ne} %z } => { %r = R{e}  %zf }
pattern vjcc.zf.e .mask { body(0x40,  0x40)  ; %r = R{e}  %z } => { %r = R{e}  %zf }
pattern vjcc.zf.ne.mask { body(0x40,  0x40)  ; %r = R{ne} %z } => { %r = R{ne} %zf }

; ── SF ── mask 0x80, idx 4
pattern vjcc.sf.ns.zero { body(0x80,  0x0)   ; %r = R{e}  %z } => { %r = R{ns} %sf }
pattern vjcc.sf.s .zero { body(0x80,  0x0)   ; %r = R{ne} %z } => { %r = R{s}  %sf }
pattern vjcc.sf.s .mask { body(0x80,  0x80)  ; %r = R{e}  %z } => { %r = R{s}  %sf }
pattern vjcc.sf.ns.mask { body(0x80,  0x80)  ; %r = R{ne} %z } => { %r = R{ns} %sf }

; ── OF ── mask 0x800, idx 5
pattern vjcc.of.no.zero { body(0x800, 0x0)   ; %r = R{e}  %z } => { %r = R{no} %of }
pattern vjcc.of.o .zero { body(0x800, 0x0)   ; %r = R{ne} %z } => { %r = R{o}  %of }
pattern vjcc.of.o .mask { body(0x800, 0x800) ; %r = R{e}  %z } => { %r = R{o}  %of }
pattern vjcc.of.no.mask { body(0x800, 0x800) ; %r = R{ne} %z } => { %r = R{no} %of }
template vjcc<FLAG, MASK, IDX, CC_SET, CC_CLEAR> {
  %rf = X86ReadFlags %f[0..5]
  %w  = launder %rf
  %m  = And %w, imm MASK
  %s  = Sub %m, imm SUB          where SUB ∈ { 0, MASK }
  %z  = X86Flag.ZF %s
  %r  = R{KIND} %z               where KIND ∈ { e, ne }
} => {
  %r  = R{ (SUB == MASK) ⊕ (KIND == ne) ? CC_SET : CC_CLEAR } %f[IDX]
}

instantiate vjcc<CF, 0x1,   0, b, ae>
instantiate vjcc<PF, 0x4,   1, p, np>
instantiate vjcc<ZF, 0x40,  3, e, ne>
instantiate vjcc<SF, 0x80,  4, s, ns>
instantiate vjcc<OF, 0x800, 5, o, no>
Kitchen table: A normal CPU has a few light-up flags after a compare: “was it smaller? was it equal? did it overflow?” A VM that wants to hide a branch will unscrew those lights, test them with a little ritual (mask, subtract, check zero), and only then jump. The table above is the ritual decoded. Once you recognize the ritual, you screw the original lights back in and write a normal if.
For operators: Read the template as a compiler rewrite, not as documentation of Intel flags. X86ReadFlags %f[0..5] materializes CF, PF, AF, ZF, SF, OF in that index space; AF is unused here. launder kills any TBAA / flags-alias assumption so the And/Sub cannot be thought of as still being EFLAGS. SUB is either 0 or MASK; KIND is either e or ne. The xor in the replacement opcode is the truth table: (subtracted the mask) XOR (tested ne) selects CC_SET versus CC_CLEAR. Instantiations skip AF (idx 2) because Tencent’s published VJCC family does not virtualize JA/JBE-style AF combinations; CF, PF, ZF, SF, OF cover the Jcc set they observed. If a later Tencent build adds a combined CF&ZF predicate as its own handler chain, this DAG table will miss and the lifter will halt on a symbolic destination — which is exactly the “optimizations incomplete” fork the authors describe.

A worked CF example, staying inside the published rules. body(0x1, 0x0) plus R{e} on ZF of the subtraction is jae/jnb (ae). The same body with R{ne} is jb/jc. Switching SUB to the mask 0x1 flips the sense. That is why four patterns exist per flag rather than two. An incomplete matcher that only handles the zero-SUB forms will invert half of the recovered branches.

Simple MBA identity rule reduction

Mixed Boolean-arithmetic obfuscation rewrites a simple operation such as A + B into a nest of ANDs, ORs, XORs, additions and negations that still compute A + B. Strong MBA is ugly: high-degree polynomials over bits, constants inside bitwise operands, and identities that need 1-bit folding (MBA-Blast, SiMBA) or synthesis. Tencent VM, in Aftermath Labs’ telling, uses trivial identities applied recursively. The list they publish is exhaustive for this VM. Put the identities in an inst-combine ruleset, run to a fixed point, and the MBA layer falls over.

(A|B) + (A&B)        = A + B
(A^B) + (A&B)        = A | B          [from ((B^A)+(B&A)) → x|y]
(A|B) - (A&B)        = A ^ B
((A|B)^A) + A        = A | B          [also A + ((A|B)^A)]
~((~A ^ ~B) | ~A)    = A & B          [De Morgan variant]
~(((A^B) & ~B) ^ ~A) = A & B
A - (A - (A&B))      = A & B
B - (((A&B)&B) ^ B)  = A & B

((c&A) ^ A) | A      = A              [+ all operand orderings]
(A & ((A^c) | c))    = A
(((A&c) ^ A) | A)    = A
(((c|A) ^ c) | A)    = A
(((A^B)|B) & A) + B  = A + B          [((A^B)|B)=A|B, (A|B)&A=A]
(A&B) + (A|B)        = A + B

~(A - c)             = -A + (c-1)     [reported as -A, const folded]
~( <A+c gadget> )    = -A + c
Worked rewrite of Tencent MBA identities including OR-AND addition
One identity from the published list, with the surrounding family. Recursion is the only amplifier; fixed-point inst-combine is the solvent. Additional explainer produced for this draft.
Kitchen table: MBA is writing “five” as “the number of fingers on one hand plus the number of fingers that are not thumbs, minus the thumbs that are also fingers.” It is still five. Tencent’s version of this joke uses only a handful of punchlines, repeated. If you keep applying the punchlines backwards, the joke becomes the number again. Hard MBA invents new punchlines faster than you can catalogue them. This VM did not.
For operators: These identities are the linear, local, bitwidth-agnostic subset. (A|B)+(A&B)=A+B is the classic bitwise-add split (OR is the sum without carry, AND is the carry). (A|B)-(A&B)=A^B is the same split with a subtract. The four “= A” absorbing forms are junk-code kill: they exist to bulk the expression without changing the value. ~(A-c)=-A+(c-1) is two’s-complement identity; Aftermath Labs note it is reported as -A after constant folding. If you already run a real inst-combine, you may not even need that last row as a special case. What you must not do is reach for a heavy MBA-Blast pass as the first tool: it will work, and it will hide the fact that a twelve-line ruleset was enough. Keep the cheap rules, measure the residual, only then escalate. The 2026 A2-MBA line of work (architectural hardening, anti-generalization) is the vendor-side answer to exactly this cheap-ruleset failure; Tencent VM, as published, has not taken that step.

A second worked identity, written as bits

Take A – (A – (A & B)). Inner A – (A & B) clears from A every bit that is set in both A and B, which is A & ~B. Subtracting that from A again yields A – (A & ~B) = A & B. The VM can nest this inside an OR-identity and an absorbing form and the SSA tree looks busy. After two rewrites it is a bitwise AND. That is the entire MBA story on this target: depth, not novelty.

Lowering

Before the SSA IR can become native code again, the authors concretize RSP to an arbitrary constant at the start of lifting. Existing optimizations then fold RSP modifications. Concrete RSP is also the heuristic for “are we at a VMEXIT?” They note you can keep RSP symbolic; this is simply the approach they already used on Themida, VMProtect, vxlang, Tencent VM and the Denuvo anti-cheat driver.

Concrete RSP rewrite — original looping diagram of folding VM stack traffic against a sentinel RSP. Original looping diagram. Source: original article.

The original stack-frame size still has to be recovered. Dynamic stack allocation is rare in AMD64 PE functions, so at sink points — calls and boxed instructions — RSP reveals the original frame. Aftermath Labs call this the stack-frame high-water mark.

Stack-frame high-water mark — original looping diagram of measuring guest RSP at boxed and CALL sinks. Original looping diagram. Source: original article.

Once the high-water mark is known, the recompiler rebuilds a prolog and epilog that represent the original frame. If the recompiled code spills, the original frame is placed after the recompiler’s own frame. Any reference to RSP at or above the return address is adjusted by the new spill size. For functions without dynamic stack allocation, that is a correct devirtualized frame.

Spill-frame layout — original looping diagram of placing the original guest frame after recompiler spill space. Original looping diagram. Source: original article.
Kitchen table: Lowering is putting the translated recipe back on a card that a normal cook can hold. You have to remember how much counter space the original dish used, even though your translator also needed a little extra counter for scratch paper. If you mix the two counters, the cook knocks the flour on the floor. The high-water mark is the stain the original cook left on the original counter. Measure the stain, then put your scratch paper beside it, not on top of it.
For operators: Implementation notes that follow from the source without adding unpublished math. (1) The sentinel RSP value must be far from the image base and far from .tvm0, or constant folding will confuse stack slots with bytecode. (2) High-water mark is a max, not a last: a function that CALLs from two different depths must keep the deeper one. (3) Spill adjustment is a relocation of every recovered memory operand whose SSA origin is “RSP + k with k >= return-address slot.” Miss one and you have a silent stack-smash in the recompiled function. (4) Functions that do alloca, or that probe a huge frame with chkstk in a way the boxed-sink heuristic never sees, are in the uncovered tail. That is a plausible contributor to ACE-CORE.sys sitting at 89.8 percent while ACE-GAME.sys sits at 99.3.

Deobfuscation coverage statistics

Coverage is measured per driver as the fraction of virtualized functions successfully recompiled to native code. A virtualized function is identified by its entry trampoline (E9 rel32 padded with int3). After devirtualization the trampoline retargets from the .tvm0 bytecode interpreter to the recompiled .devirt code. Functions the tool could not recover still point into .tvm0.

DriverVirtualized functionsDevirtualizedCoverage
ACE-GAME.sys13413399.3 %
ACE-BASE.sys32931395.1 %
ACE-BOOT.sys23521993.2 %
ACE-CORE.sys16715089.8 %
Total86581594.2 %
Per-driver static recovery published by Aftermath Labs. Source: original article.

Across the four kernel drivers, 815 of 865 virtualized functions (94.2 %) were fully recovered to native code.

Bar chart of recovery coverage on four ACE kernel drivers
Same numbers as the original table, drawn as bars so the ACE-CORE.sys dip is visible at a glance. Additional explainer produced for this draft.
Decompiled Devirt_DriverEntry of ACE-GAME.sys after static recovery
Devirtualized driver entry of ACE-GAME.sys. Aftermath Labs say most recovered functions are this clean. Source: original article.

The recovered DriverEntry is ordinary WDM: it fills MajorFunction slots for create/close, device control and shutdown, sets DriverUnload, calls AceGameStartup, then RtlInitUnicodeString / IoCreateDevice / IoRegisterShutdownNotification on \Device\ACE-GAME. That is the point of the screenshot. After a 94 percent recovery you are no longer staring at a dispatcher. You are staring at a driver.

Kitchen table: Coverage is a school test. 815 out of 865 questions were answered in readable English. Fifty questions are still in the private language. ACE-GAME.sys almost aced the test. ACE-CORE.sys left more blank. For a working analyst that is the difference between “I can read this module” and “I can read most of this module and I know which stubs still go into the interpreter.”
For operators: How to reproduce the census without their toolchain: scan for E9 rel32 whose remaining bytes to the next boundary are CC, whose target lands in a section named .tvm / .tvm0 / similar, and whose function start is referenced like a normal entry. After a rewrite pass, the same sites should either still point into .tvm0 (failure) or point at a new .devirt range (success). Do not count ordinary E9s in .text. Daax being given the files is the source’s answer to “trust the 94.2 percent.” If you need a lab-internal confirmation, that is the third-party path they already named.

Limitations

Ranged for-loops collide with the aggressive indirect-control-flow optimizations. In a construct such as for (int i = 0; i < 10; i++), the virtualized i < 10 becomes a VJCC. Because i starts at 0, constant folding keeps only the loop-taken destination, which is the back-edge. The lifter has already lifted that back-edge, so it stops and the recovered function is an infinite loop. Aftermath Labs say relifting back-edges may be required, symbolizing everything except VIP.

Kitchen table: If a translator decides on page one that a recipe says “stir forever” because the first reading of the counter was zero, they will never see the line that says “stop at ten.” The fix is to treat the counter as unknown and keep the “stop at ten” door in the text, even though today’s shopping list starts at zero.
For operators: This is the classic “over-concretization of a loop guard” bug, cousin to under-unrolling in mixed static-dynamic lifters. A practical patch consistent with the source: when a VJCC DAG matches and one destination is a seen VIP, still emit both successors if the predicate SSA still depends on a value that is not VIP and not a bytecode constant. In other words, constant-fold bytecode, do not constant-fold the guest induction variable. The authors’ suggested relift (symbolize everything except VIP) is the conservative version of the same rule.

Where this sits next to Themida, VMProtect and Denuvo

Aftermath Labs are explicit that the same lowering tricks were already used on Themida, VMProtect, vxlang, Tencent VM and the Denuvo anti-cheat driver. Their May 2026 Themida article argued the techniques apply to pretty much every virtual-machine obfuscator with minor modifications. The Tencent article is the kernel-anti-cheat data point in that series.

What repeats across the family is the silhouette: GPR save into a context, a shared dispatcher, bytecode-chosen handlers, boxed native escapes, MBA on the ALU, virtualized flags, a VIP that can be named. What changes is the compatibility skin. Tencent’s skin is CET (RDSSPQ / INCSSPQ 2) and a phantom-unwind SEH story. VMProtect’s skin is different handler shuffling and mutation. Themida’s skin is a different bytecode encryption and a different boxed set. A lifter that is built as a framework with per-VM guidance plugins will keep winning. A lifter that is a single hardcoded Tencent script will die on the next vendor.

For operators: Their own product, CodeDefender, is pitched at the end of the source as the thing that tries to hinder lifting-based attacks and guided symbolic evaluation. Read that as the authors telling you where they think the real obfuscation problem has moved: not “more handlers,” but anti-lifting architecture. The 2026 MBA hardening literature (architectural state bridging, anti-generalization against rule explosion) is the academic twin of that pitch. If you are buying a protector because it says “virtual machine” on the box, this article is the reason that sentence is no longer enough.

Detection, hunting, and what this looks like on disk

Defenders do not need to recover 815 functions to benefit from the write-up. The published tells are already enough to build a hunting checklist for VM-protected kernel modules and for user-mode cousins that share the silhouette.

Static tells

  • PE sections named .tvm, .tvm0, or obvious siblings, with a high entropy data stream and a small native dispatcher.
  • Function entries that are a single E9 rel32 followed by int3 padding, targeting that section.
  • A language-specific exception handler covering an unusually large .pdata range that includes a dispatcher.
  • RDSSPQ (hint-NOP space) followed by a conditional INCSSPQ with immediate-class 2, inside kernel code that is not the Windows CET runtime itself.
  • A dummy function that is never called, whose unwind codes describe nine non-volatile saves, and whose address is stored at [RBP+0] from a VM entry that uses UWOP_SET_FPREG.

Behavioral tells

  • A kernel module that spends a large fraction of sampled IP in one dispatcher loop with a rolling bytecode pointer.
  • Frequent native context restore / execute / recapture around CPUID, MSR, or similarly unmodelable instructions (the boxed path).
  • CALL sites that never RET, paired with shadow-stack pointer adjustment on CET-enabled hosts.

ATT&CK and CWE, kept honest

LensIDHow it applies here
Protector as defense evasionT1027.002 Software PackingSelected functions are replaced by an interpreter and bytecode. This is packing, even when the vendor is anti-cheat rather than malware.
Protector as obfuscated fileT1027 Obfuscated Files or InformationMBA identities and virtualized flags hide the original ALU and branches from linear disassembly.
Analyst technique, not attackerT1127 / researchGuided symbolic execution is the analyst’s tool. Do not file it as an adversary technique just because the subject is anti-cheat.
If an ACE driver is later buggyCWE-269 / CWE-787 etc.Not claimed by this source. Keep LPE hunts in a different ticket from packing hunts.
Exception-path correctnessCWE-755 Improper Handling of Exceptional ConditionsThe phantom-unwind design is how Tencent avoids this class inside the VM. A broken port of the same idea would land here.
Mapping for hunters and program owners. The source is a protector analysis, not a CVE.
Kitchen table: Hunting this without being a reverse engineer: ask your software-inventory people whether a kernel driver has a section whose name looks like a TV station (.tvm) and whose entry points are one-instruction jumps into that section. That question already finds the population. The rest of this article is what a specialist does once the population is on a desk.

What this article does not give you

  • A ready-to-run ACE unpacker, a recovered ACE-*.sys image, or a cheat.
  • A filled-in Tencent bytecode decoder beyond the heuristics and IR already published by Aftermath Labs.
  • A claim that every in-house VM is this weak. The authors’ point is that classic VMs of this silhouette are. Their CodeDefender pitch is that a protector can be built to fight this exact pipeline.
  • A claim of a Microsoft-serviced vulnerability. CET and SEH compatibility here are engineering, not CVEs.

Key Takeaways

  • Tencent VM is a classic GPR-save / shared-dispatcher / bytecode-handler virtual machine sitting in ACE kernel drivers, with boxed native escapes for the rest of AMD64.
  • CET compatibility is a runtime feature check (RDSSPQ) plus INCSSPQ 2, because CALLs into a dispatcher that never returns would otherwise exhaust the shadow stack.
  • SEH compatibility is a single large .pdata blanket, a frozen RBP frame pointer, nine staged non-volatiles, a dummy function with phantom unwind info at [RBP+0], and ExceptionContinueSearch.
  • Guided symbolic execution inlines the dispatcher (CALL+int3 heuristic), promotes .tvm0 bytecode loads, tracks VIP from the last .tvm0 pointer stored into the VM context, and refuses to unroll seen VIPs.
  • Virtual Jcc is a four-pattern-per-flag DAG on CF/PF/ZF/SF/OF. The published template is enough to rewrite it back to native Jcc and to extract both targets.
  • The MBA layer is a closed list of trivial identities. Inst-combine to a fixed point is the published solvent; this is not a job that, on this VM, requires MBA-Blast as the first tool.
  • Lowering reuses a concretized RSP, a stack-frame high-water mark at boxed/CALL sinks, and a spill-then-guest frame layout. Reported coverage is 815/865 functions (94.2 percent) across four ACE drivers, with ACE-GAME.sys at 99.3 percent.
  • The remaining failure mode they highlight is ranged for-loops whose VJCC is over-folded into a single back-edge. Relift with everything except VIP symbolic is the stated way out.

Defensive Recommendations

  1. Inventory VM-protected kernel modules. Section names, E9+int3 trampolines into those sections, and oversized .pdata ranges are enough for a first census of ACE and of look-alikes.
  2. Do not treat “it is virtualized” as the end of an IR ticket. If a kernel driver in your estate is VM-protected, assume a capable shop can recover the majority of it statically, because that is what this paper just showed on ACE.
  3. Hunt RDSSPQ/INCSSPQ pairings in third-party kernel code. They are a CET-hygiene tell for interpreters that CALL and never RET. Baseline the Windows CET runtime so you do not page yourself.
  4. If you ship a protector, stop shipping this silhouette. A shared dispatcher, a context-resident VIP, boxed sinks, and recursive trivial MBA are a solved combination. Anti-lifting design (the authors’ CodeDefender claim, and the 2026 MBA-hardening literature) is the actual product conversation.
  5. If you ship an anti-cheat driver, budget for interoperability analysis. Aftermath Labs’ stated purpose is Linux/Proton interoperability. Closed kernel ACE on titles that players reasonably run under Proton will be reverse engineered. Plan a supported path instead of relying on the VM to be the path.
  6. Keep recovered artefacts inside the legal envelope the source used. Interoperability research and independent validation with a named third party is one thing. Redistributing unpacked ACE images or a cheat kit is another.
  7. For malware crews that reuse the same VM grammar: the VJCC DAGs and MBA identities in this article are reusable detections of the grammar, not of ACE specifically. YARA on RDSSPQ+INCSSPQ+dispatcher loops will overfit; prefer structural rules in your lifter and a trampoline census on disk.
  8. Watch ACE-CORE.sys-class tails. A 10 percent uncovered set is where interesting control flow hides. If you are doing a security review of a VM-protected driver, start with the functions that still point into .tvm0 after a recovery pass.

Conclusion

Aftermath Labs close where they opened. Classic virtual-machine obfuscation is not strong in the face of guided symbolic evaluation. Simple compiler passes — constant promotion, constant folding, and a short MBA identity list — were enough to simplify Tencent VM. They present the ACE numbers as evidence, the CET and SEH chapters as proof that even a VM which did its Windows-hygiene homework still lifts, and their CodeDefender product as the place they have tried to block this exact attack. Analyzing hardened targets, they write, is how a protector learns which pitfalls are obvious only after someone has driven a lifter through them. Consulting and product contact sit on the Aftermath Labs site; the post’s visible mailbox is Cloudflare-protected.

For the rest of us the lesson is plainer. A virtual machine that saves GPRs, dispatches on bytecode, boxes the leftover ISA, and obfuscates arithmetic with a pocket sheet of identities is a compiler problem. Compiler problems get solved. The interesting engineering in this paper is not a magic deobfuscator. It is the discipline of naming VIP, naming the boxed sinks, naming the Jcc DAGs, and then letting an ordinary optimizer finish the job.

Kitchen table: If a lock is just a translator between a public language and a private one, a better translator opens it. Tencent VM, as published, is that kind of lock. The next generation of protectors will have to stop being translators and start being things a compiler cannot name. Until they do, articles like this one will keep turning private kitchens back into readable cards.

Original text: “Static Devirtualization of Tencent VM” by author not clearly listed (site: Aftermath Labs) at Aftermath Labs.

oxfemale Vulnerability research, reverse engineering, and exploit development.
// Discussion