13 Commits
Author SHA1 Message Date
ApfelTeeSaft 8242ffb49b Merge pull request #13 from KallDrexx/loopback_instruction_fix
Loopback jumps should occur on the same cpu address of the jump
2026-01-17 17:55:48 +01:00
KallDrexx e76cec3d6d Loopback jumps should occur on the same cpu address of the jump
When a function is decompiled in the middle of a function, and that function
has a jump point prior to the function's entry point, we need to add a
fake virtual instruction that jumps back to the start of the function.

The virtual instruction needs to be the last instruction in the
ordered instruction set.

We previously accomplished that by adding an instruction at the
location of the function entrypiont minus one. However, this was
failing in cases where a system would cause an interrupt right on
the virtual address. Emulators would then save the virtual
instruction's address to the stack, and jump back into that address
once RTI occurs.

This fails because the virtual address can't be decompiled, because
legit code doesn't exist at that address.

To fix this, I updated the `SubAddressOrder` property to allow for negative
values. This allows the loopback instruction to be on the correct CPUAddress
while still being ordered as expected.

Also ensured that virtual addresses do not get labels, as they are not
actually valid jump targets.
2026-01-17 11:44:40 -05:00
ApfelTeeSaft 26af1f4aa6 close enough ig 2025-10-23 11:34:11 +02:00
Matthew Shapiro bc3da6b960 Decompilation should be possible from the first byte of a code region (#12) 2025-10-17 20:30:54 -04:00
Matthew Shapiro c71aa4d6b2 Consider an invalid instruction the end of a function trace (#11)
When tracing a function, previously we were throwing an exception
if we encountered an invalid / unknown op code. This is needed because
we don't know how many bytes the instruction contains and thus can't
accurately predict where the next instruction would be.

However, some roms (like super mario bros) use an always taken
branch instruction to save bytes instead of a jump. In this case
the next byte after the branch is an invalid op code that will never
actually be hit.

So this change makes it so that invalid operations act as an end
of function markers, but throw a warning in the console. That
allows unconditional/always branches to work, and for games that
actually use these unofficial op codes they are able to get hints
in the debug window.
2025-10-12 14:31:27 -04:00
Matthew Shapiro bcd9ed4b87 Merge pull request #10 from KallDrexx/support_virtual_instructions
Allow for sub address instructions.



When a wrap around scenario is detected during single function tracing, if the address prior to the "entry point" is a single byte, then we do not have any space to add the required jump call.

This fixes that by adding the concept of sub address instructions. This allows adding instructions at runtime that get sorted correctly against the real instructions from the ROM.

This not only solves the wrapping issue, but also allows for adding hooks at runtime.
2025-10-11 23:17:12 -04:00
KallDrexx c0380bd198 Allow for sub address instructions.
When a wrap around scenario is detected during single function tracing,
if the address prior to the "entry point" is a single byte, then we do
not have any space to add the required jump call.

This fixes that by adding the concept of sub address instructions. This
allows adding instructions at runtime that get sorted correctly against
the real instructions from the ROM.

This not only solves the wrapping issue, but also allows for adding hooks
at runtime.
2025-10-11 23:11:37 -04:00
Matthew Shapiro c05b7199c9 Merge pull request #9 from KallDrexx/single_function_decompile_wraparound
Fix wraparound bug
2025-10-11 21:10:03 -04:00
KallDrexx ab9a1fa313 Fix wraparound bug
Since the entry point for analysis could be in the middle of a loop,
we need to guarantee that a jump is dedicated to the entrypoint, so
that an instruction that comes before the "entrypoint" will redirect
back to the entrypoint after execution
2025-10-11 21:00:22 -04:00
Matthew Shapiro bbcd2a6aad Merge pull request #8 from KallDrexx/single_function_decompile
Add code path to decompile/disassemble a single function
2025-10-11 16:03:41 -04:00
KallDrexx 826747aacb Fixed incorrect ordering of instructions 2025-10-11 15:49:10 -04:00
KallDrexx 8e811e2cbc Some fixes 2025-10-10 23:26:38 -04:00
KallDrexx 6d3ec6c2c8 Initial single function decompiler implementation 2025-10-10 23:06:03 -04:00
5 changed files with 327 additions and 1 deletions
+73
View File
@@ -0,0 +1,73 @@
# Changelog
All notable changes to this project will be documented in this file.
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/).
---
## [v1.1.1] - 2025-10-18
### Added
- Decompilation can now start from the **first byte of a code region** (#12).
### Changed
- Invalid instructions are now considered the **end of a function trace** (#11).
- Improved stability of function tracing and edge-case instruction handling.
### Fixed
- Minor internal decompiler logic bugs.
- General performance and reliability improvements across builds.
### Notes
- All Windows builds now include both GUI and CLI versions.
- Linux and macOS builds are CLI-only.
- All binaries are **self-contained** and **do not require a .NET runtime**.
**Contributors:**
[@KallDrexx](https://github.com/KallDrexx)
---
## [v1.1.0] - 2025-10-12
### Added
- Implemented **single function decompiler** with proper tracing.
- Added support for **sub-address instructions** and **virtual instructions** (#10).
- Added **CI/CD workflow** (`build-release.yml`) for automated builds and packaging.
- Added **multiple platform releases**:
- Windows (x64, x86, ARM64) with GUI + CLI
- Linux (x64, ARM64) CLI
- macOS (x64, ARM64) CLI
### Changed
- Improved function discovery to handle **wraparound and disassembly boundaries** (#9).
- Reworked `ToString()` formatting for instructions for clarity.
- Improved tracing logic for **unreferenced instruction analysis** (#6).
- Decompiler now directly jumps to instructions that appear within other instructions.
### Fixed
- Fixed incorrect ordering of instructions in output.
- Fixed 16KB ROMs not decompiling (#7).
- Fixed nullability warnings.
- Fixed various stability issues in the decompiler core.
**Contributors:**
[@ApfelTeeSaft](https://github.com/ApfelTeeSaft), [@KallDrexx](https://github.com/KallDrexx)
---
## [v1.0.0] - 2025-05-15
### Added
- **Initial release** of the NES Decompiler.
- Included both **CLI** and **GUI** builds for Windows (x64).
- Added base decompilation engine and ROM handling logic.
- Added initial README and documentation.
**Contributors:**
[@ApfelTeeSaft](https://github.com/ApfelTeeSaft)
---
## [Unreleased]
- Planned improvements to function boundary detection.
- Optimizations for recursive instruction analysis.
- Additional architecture support under evaluation.
@@ -0,0 +1,8 @@
namespace NESDecompiler.Core.Decompilation;
/// <summary>
/// A set of code that may contain executable code
/// </summary>
/// <param name="BaseAddress">Where the first byte of the region can be found from the CPU's memory map</param>
/// <param name="Bytes">The set of data to pull code out of</param>
public record CodeRegion(ushort BaseAddress, ReadOnlyMemory<byte> Bytes);
@@ -0,0 +1,67 @@
using NESDecompiler.Core.Disassembly;
namespace NESDecompiler.Core.Decompilation;
/// <summary>
/// Represents an independently decompiled function
/// </summary>
public class DecompiledFunction
{
/// <summary>
/// The CPU address where the address' first instruction is located
/// </summary>
public ushort Address { get; }
/// <summary>
/// The instructions that make up this function in the correct order in which they should be
/// executed.
/// </summary>
public IReadOnlyList<DisassembledInstruction> OrderedInstructions { get; }
/// <summary>
/// Location and the labels of all jump and branch targets within this function
/// </summary>
public IReadOnlyDictionary<ushort, string> JumpTargets { get; }
public DecompiledFunction(
ushort address,
IReadOnlyList<DisassembledInstruction> instructions,
IReadOnlySet<ushort> jumpTargets)
{
Address = address;
JumpTargets = instructions
.Where(x => jumpTargets.Contains(x.CPUAddress))
.Where(x => x.Label != null)
.Where(x => x.SubAddressOrder == 0) // only real instructions should be jumped to
.ToDictionary(x => x.CPUAddress, x => x.Label!);
// We need to order the instructions so that the starting instruction is the first one encountered.
// We can't just rely on the CPU address, because a function may jump to a code point earlier than
// the first instruction.
var entryPointInstructions = instructions.Where(x => x.CPUAddress == address)
.Where(x => x.SubAddressOrder >= 0);
var initialInstructions = instructions
.Where(x => x.CPUAddress > address)
.OrderBy(x => x.CPUAddress)
.ThenBy(x => x.SubAddressOrder);
var trailingInstructions = instructions
.Where(x => x.CPUAddress < address)
.OrderBy(x => x.CPUAddress)
.ThenBy(x => x.SubAddressOrder); // real instructions before virtual ones
// If there was a loopback jump point at the function address, put that here. This is required
// because if an emulator is executing a virtual loopback instruction and an IRQ occurs, this
// will cause the virtual instruction to be saved to the stack, and that can cause the entry
// point to be wrong.
var loopbackInstructions = instructions.Where(x => x.CPUAddress == address)
.Where(x => x.SubAddressOrder < 0);
OrderedInstructions = entryPointInstructions
.Concat(initialInstructions)
.Concat(trailingInstructions)
.Concat(loopbackInstructions)
.ToArray();
}
}
@@ -0,0 +1,170 @@
using NESDecompiler.Core.CPU;
using NESDecompiler.Core.Disassembly;
namespace NESDecompiler.Core.Decompilation;
public static class FunctionDecompiler
{
/// <summary>
/// Traces and decompiles a single function
/// </summary>
/// <param name="functionAddress">The CPU address of the entry point of the function to decompile</param>
/// <param name="codeRegions">All available regions of bytes that could contain instructions for the function</param>
public static DecompiledFunction Decompile(ushort functionAddress, IReadOnlyList<CodeRegion> codeRegions)
{
var instructions = new List<DisassembledInstruction>();
var jumpAddresses = new HashSet<ushort>();
var seenInstructions = new HashSet<ushort>();
var addressQueue = new Queue<ushort>([functionAddress]);
while (addressQueue.TryDequeue(out var nextAddress))
{
if (!seenInstructions.Add(nextAddress))
{
if (nextAddress == functionAddress)
{
// This means a branch occurred that caused the flow to wrap around to instructions preceding
// the function entrance. This usually happens when there is a jump/branch to right before the
// entrypoint, usually due to decompiling in the middle of a loop. To fix this, we need to add
// a jump back to the function entrypoint.
if (functionAddress == 0x00)
{
const string message = "Wrap around instruction detected for a function at 0000, but that " +
"doesn't make sense";
throw new InvalidOperationException(message);
}
var addressHigh = (functionAddress & 0xFF00) >> 8;
var addressLow = functionAddress & 0x00FF;
var jumpInstruction = new DisassembledInstruction
{
Info = InstructionSet.GetInstruction(0x4C),
CPUAddress = nextAddress,
Bytes = [0x4C, (byte)addressLow, (byte)addressHigh],
TargetAddress = functionAddress,
// Make sure they appear before the function address
SubAddressOrder = -1,
};
instructions.Add(jumpInstruction);
}
continue;
}
var instruction = GetNextInstruction(nextAddress, codeRegions);
if (instruction == null)
{
// Consider no instruction the end of the function. This is usually the case
// with an always taken branch
continue;
}
instructions.Add(instruction);
// Ensure the function entrypoint has a label
if (instruction.CPUAddress == functionAddress && instruction.Label == null)
{
instruction.Label = $"sub_{functionAddress:X4}";
jumpAddresses.Add(functionAddress);
}
if (IsEndOfFunction(instruction))
{
continue;
}
if (instruction.TargetAddress != null)
{
jumpAddresses.Add(instruction.TargetAddress.Value);
addressQueue.Enqueue(instruction.TargetAddress.Value);
}
if (!instruction.IsJump)
{
addressQueue.Enqueue((ushort)(nextAddress + instruction.Info.Size));
}
}
// Add labels for any jump targets
foreach (var instruction in instructions)
{
// Only real instructions should have a label, virtual ones should not
if (jumpAddresses.Contains(instruction.CPUAddress) && instruction.SubAddressOrder == 0)
{
instruction.Label = $"loc_{instruction.CPUAddress:X4}";
}
}
return new DecompiledFunction(functionAddress, instructions, jumpAddresses);
}
private static DisassembledInstruction? GetNextInstruction(ushort address, IReadOnlyList<CodeRegion> regions)
{
var relevantRegion = regions
.Where(x => x.BaseAddress <= address)
.Where(x => x.BaseAddress + x.Bytes.Length > address)
.FirstOrDefault();
if (relevantRegion == null)
{
var message = $"No code region contained the address 0x{address:X4}";
throw new InvalidOperationException(message);
}
var offset = address - relevantRegion.BaseAddress;
var bytes = relevantRegion.Bytes.Span[offset..];
var info = InstructionSet.GetInstruction(bytes[0]);
if (!info.IsValid)
{
var message = $"Warning: encountered unknown op code 0x{bytes[0]:X2} at address 0x{address:X4}";
Console.WriteLine(message);
return null;
}
if (bytes.Length < info.Size)
{
var message = $"Opcode {info.Mnemonic} at address 0x{address:X4} requires {info.Size} bytes, but only " +
$"{bytes.Length} are available";
throw new InvalidOperationException(message);
}
var instruction = new DisassembledInstruction
{
Address = (ushort)offset,
CPUAddress = address,
Info = info,
Bytes = bytes[..info.Size].ToArray(),
};
Disassembler.CalculateTargetAddress(instruction);
return instruction;
}
private static bool IsEndOfFunction(DisassembledInstruction instruction)
{
// RTI and RTS are obviously the end of a function. We consider BRK and JSR
// to be the end of a function as well because an RTI or RTS will do a function
// call into the next instruction. This is required because RTI/RTS could be
// returning based on a modified stack, and therefore we are not guaranteed to
// be returning to the expected spot.
if (instruction.Info.Mnemonic is "JSR" or "BRK" or "RTI" or "RTS")
{
return true;
}
// Since we don't know where we are jumping at compile time, this will be treated
// as a function call, thus we consider it the end of the function.
if (instruction.Info.AddressingMode == AddressingMode.Indirect)
{
return true;
}
return false;
}
}
@@ -72,6 +72,14 @@ namespace NESDecompiler.Core.Disassembly
/// </summary>
public bool IsJump => Info.Mnemonic == "JMP" || Info.Mnemonic == "JSR";
/// <summary>
/// Determines the order of this instruction within a single address space. This is mostly
/// needed in the cases that additional instructions are needed to be added in the same
/// address location at runtime. Can be used to add runtime hooks or to work around
/// decompilation issues. Should be 0 for all native instructions from a ROM.
/// </summary>
public sbyte SubAddressOrder { get; set; }
/// <summary>
/// Returns a string representation of this instruction
/// </summary>
@@ -439,7 +447,7 @@ namespace NESDecompiler.Core.Disassembly
/// Calculates the target address for branch and jump instructions
/// </summary>
/// <param name="instruction">The instruction to process</param>
private void CalculateTargetAddress(DisassembledInstruction instruction)
public static void CalculateTargetAddress(DisassembledInstruction instruction)
{
if (instruction.Info.AddressingMode == AddressingMode.Relative)
{