Martin's Blog

A Safer Wrapper Is Not a Modernisation Strategy

Moving a legacy product to a modern operating system is useful engineering. It can deliver a maintained kernel, current libraries, better compiler defaults, ASLR, mandatory access control, namespaces, cgroups and service supervision. All of those can make exploitation harder or contain a failure.

It can also leave the product’s old trust model almost entirely intact.

The new kernel may still launch the same privileged packet engine, the same configuration daemon, the same PHP application and the same Perl and shell helpers. A process running as root retains broad authority unless capabilities, namespaces or mandatory policy materially constrain it. A shell command built from attacker-controlled text remains a shell command. A parser written in C remains memory unsafe even when systemd starts it.

The wrapper buys time and lowers some risks. It does not retire the engineering debt inside it.

NetScaler 15.1 makes the distinction visible

I compared NetScaler VPX 15.1 preview build 9.33 with the 14.1 appliance from an earlier assessment. The platform had moved from FreeBSD 11.4 and UFS to Linux 6.6, systemd, ext4, cgroup v2, Linux network namespaces and a DPDK/VFIO data plane. Several old vulnerability paths were also fixed. That is substantial engineering, not a cosmetic port.

The application architecture showed more continuity. Apache and PHP still served management; the appliance bundled Perl 5.30 and 172 Perl scripts for tasks including backup, support collection, monitoring and CallHome. The packet engine, configuration daemon, authentication daemon and most proprietary services remained native programs running as root inside the same blx network namespace.

NetScaler VPX 14.1 and 15.1 preview architecture comparison showing a Linux platform rebase around recognisable packet, control, web and scripting components

Observed platform change and component continuity in the locally assessed builds. This does not establish identical internals or a clean-sheet rewrite.

Some of the new Linux controls were mechanisms waiting for policy:

Control in the examined previewWhat was observedSecurity consequence
Network namespacesA sparse root network namespace and a separate blx network namespace containing most NetScaler servicesNetworking resources were separated, but the boundary did not itself isolate filesystems, process memory or user identities; most NetScaler services also shared blx
SELinuxInfrastructure present; the kernel booted with enforcing=0, and runtime AVC records showed permissive=1Policy violations were logged but not blocked in the observed boot
cgroup v2MountedI did not establish whether effective per-service CPU, memory or process limits were configured
systemdExplicit service lifecycle and restart policyUseful operational control; the examined CallHome unit had no meaningful sandbox directives
PIE and ASLRBroader PIE coverage for control daemonsHigher exploitation cost for those programs; the root packet engine remained a non-PIE executable
Service identitiesSome base services used dedicated usersMost proprietary services still ran as root

These observations apply to one preview build. They illustrate the difference between installing a control and making it carry a security boundary.

CVE-2026-88771 makes the point. In the examined 14.1 build, a root Perl script passed attacker-shaped login text to a shell. Build 15.1-9.33 fixed the path with strict validation, shell-free find invocation and a path check. Had the injection remained, ASLR, NX and stack canaries would not have stopped an explicitly invoked shell; SELinux was not enforcing and the service ran as root alongside important peers. cgroup limits could constrain resource abuse after execution, not prevent the command. The application fix removed the vulnerability; the platform rebase did not make it unnecessary.

A four-part modernisation programme

One practical approach is to organise modernisation around four measurable risk reductions:

  1. Move to a maintained platform and turn its security mechanisms into enforced policy.
  2. Replace dangerous scripting glue with small, defensive components.
  3. Re-architect exposed native code into memory-safe services and libraries.
  4. Prevent supported configurations from publishing administrative and unsafe interfaces through the data plane, and verify external reachability at the deployment boundary.

Delivery should overlap. Exposure reduction starts on day one because it is usually the fastest way to reduce compromise probability.

PartTypical effortTime to useful risk reductionConcrete risk reductionWhat it does not solve by itself
1. Modern platform and containmentMediumWeeks to monthsRaises exploitation cost and narrows post-compromise accessLogic flaws and overpowered application interfaces
2. Replace privileged scriptsLow to mediumDays to monthsRemoves shell interpretation and narrows privileged operationsNative memory corruption and wider trust failures
3. Memory-safe re-architectureHighMonths to yearsRemoves unsafe parsing and creates narrower service boundariesAuthorisation, protocol and privilege errors
4. Prevent unsafe exposureLow to highImmediate controls; longer product changesRemoves supported data-plane paths to management and detects external exposureAttacks through allowed services or compromised internal principals

This is not a choice between patching and rewriting. Patch exploitable defects immediately, add containment around what must remain, then remove the bug classes and trust relationships that keep producing emergencies.

1. Turn a modern operating system into a security boundary

Use a supported kernel and userland, publish their lifetime, inventory dependencies and establish repeatable updates. Build native code with applicable protections such as PIE, full RELRO, stack protection, fortified library calls, control-flow protection and ASLR.

Then do the work that feature lists tend to omit:

Each control needs an acceptance test: attempt a forbidden file read, exhaust a worker’s memory and run a test process under the parser identity. Record whether the management plane survives and which sockets, files, devices and secrets the process can reach.

Exploit mitigations raise the cost of exploitation. Privilege separation and enforced policy reduce what a successful exploit can do. A credible platform modernisation needs both.

2. Replace the dangerous scripting glue

Perl, PHP and Python are memory managed; they do not carry the same memory-corruption class as C and C++. Their common risks are injection, unsafe deserialisation, weak authorisation, path errors and secret disclosure.

Privileged scripts are often the best early rewrite targets because they are small, understandable and surprisingly powerful. Start with scripts that:

Go and Rust are good defaults for long-lived privileged components. Python can suit narrow, unprivileged orchestration with strict input and no shell. Removing the dangerous contract matters more than the language label.

Do not ask an AI coding agent to produce an ingenious translation of a clever Perl expression. Ask it to produce boring code:

CVE-2026-88771 shows the design goal. Instead of escaping text for a privileged shell command, accept a typed event with a validated identifier, open an allowlisted file through an API and return a structured result. No command string remains for an attacker to shape.

AI lowers the cost of characterising old behaviour, extracting tests and writing replacements, while still producing plausible errors. Build a golden corpus with malformed and adversarial cases, and have a human trace every privilege boundary and external call. My legacy C experiment found that agents still disagreed about scope, compatibility and “done”.

3. Re-architect exposed native code in memory-safe languages

Choose native rewrite targets by reachability and consequence, not line count. An unauthenticated root parser deserves attention before a large offline tool.

The practical route is usually incremental:

  1. Document the wire format, state machine, error behaviour and resource limits.
  2. Put a narrow interface in front of the old component.
  3. Replace parsing with a memory-safe implementation.
  4. Where containment is required, run it in a separate restricted process behind a narrow, validated IPC interface. A library inherits the privileges of its host process and does not create that boundary.
  5. Compare old and new behaviour on captured and generated traffic.
  6. Shift traffic gradually, keep a rollback path and delete the old path when compatibility is established.

Treat translation and redesign as separate acceptance gates. Establish a safe, compatible replacement behind a narrow boundary, then simplify the protocol, privilege model and design. Changing language, behaviour and architecture in one AI-assisted step makes regressions harder to attribute.

Preserve legitimate behaviour, not every legacy accident. Document intentional security changes separately, and require the replacement to reject inputs or operations that the old implementation accepted unsafely. Differential tests should identify those differences as approved security changes rather than rewarding their reintroduction.

Rust suits parsers and libraries needing tight allocation and latency control; Go suits network services that can use garbage collection. Both prevent broad memory-error classes while still permitting denial of service, cryptographic misuse and broken authorisation.

The test programme is part of the architecture, not a final project phase. It should include:

AI can help with inventory, specification, tests, bindings and repetitive porting. Independent review remains necessary. Compilation and a green unit-test suite do not prove that twenty years of implicit behaviour survived the move.

CISA and partner agencies recommend public memory-safe roadmaps. Measure them by retired exposed failure domains, not repository percentages.

4. Prevent the appliance from publishing management services

Documentation that says “do not expose the management interface” is a weak control. Appliances are installed under pressure, addresses are reused and firewall rules drift. The product should refuse supported configurations that bind administrative, diagnostic, update or legacy services to data-plane interfaces or forward data-plane traffic to management listeners.

That boundary can be enforced inside the product:

A separate management routing table and the absence of a default route reduce accidental exposure; they do not prove isolation. A specific route, upstream NAT, reverse proxy or tunnel can still publish the service. The appliance cannot control those external systems. The deployment therefore needs an independent reachability test from outside the trusted administrative boundary.

The management plane still processes hostile input from compromised administrators, devices and internal systems. It does not need direct public reachability. The testable product requirement is narrower and stronger: no supported appliance configuration should expose a management listener through the data plane, and external exposure must be controlled and verified at the deployment boundary.

Evidence from larger migrations

Google and Microsoft show two workable paths

Microsoft’s 2019 baseline was roughly 70% memory-safety CVEs; Android reported 76%. Both chose incremental migration, while Google has published the clearer outcome.

Google / AndroidMicrosoft / WindowsLesson for a legacy appliance
StrategyPrioritise memory-safe languages for new and actively developed native code; interoperate with existing C and C++Support Rust as an internal systems language and replace selected components behind established interfacesStop adding avoidable unsafe code and replace old components in risk order
Legacy codeLeave much mature code in place while new development shiftsRetain the wider Windows estate while replacing components such as GDI region handling and building firmware and driver foundationsIncremental delivery is credible, but exposed privileged parsers need explicit deadlines
EnablersInteroperability, training, build support and review practicesNative code-generation, ABI, build, compliance and servicing integrationLanguage adoption needs a production toolchain and a maintained hybrid boundary
Published resultAndroid vulnerability and delivery metricsShipped components and tooling, without a comparable public CVE trendMeasure retired failure domains and containment outcomes, not migrated lines alone

Google’s “Safe Coding” model concentrates on new and recently changed code, where vulnerabilities are densest. Google reports that memory-safety issues fell from 76% of Android vulnerabilities in 2019 to 24% in 2024 and then below 20% in 2025. It also reports about five million lines of platform Rust, 20% fewer review revisions, 25% less review time and roughly four times fewer rollbacks for medium and large Rust changes than comparable C++ changes. Its claimed 1,000-fold vulnerability-density improvement is an internal estimate based on one Rust near-miss and historical C and C++ data, not a universal benchmark. Android still adds substantial C++; the evidence supports prioritising memory-safe new development rather than claiming a complete ban.

CVE-2025-48530, an unsafe Rust overflow found before release, also shows why the wrapper matters: Google says Scudo guard pages made it non-exploitable.

By September 2026, Rust had Tier-1 internal engineering status at Microsoft, and the Microsoft engineer describing rustc_codegen_utc reported more than 100 repositories using it. Windows also shipped Rust GDI region handling behind an existing kernel interface. Check Point later reached an out-of-bounds path from low-integrity user space. A Rust bounds check produced a kernel panic instead of an unchecked access, but the panic still caused a Blue Screen. Microsoft fixed the denial-of-service path in 2025. The incident shows both outcomes: one memory- corruption path was blocked, while availability and failure handling still required design and testing.

Our roadmap combines Google’s focus on new code with Microsoft’s surgical replacement. It gives priority to old code that is unauthenticated, reachable or privileged, while adding containment and management-plane controls. Stable, unprivileged and unreachable code can wait; an internet-facing root parser cannot claim the same discount merely because it is old.

AI changes the price, not the proof

GitHub’s Copilot runtime migration is a useful process case. Coding agents helped replace the runtime with about 832,000 lines of production Rust. The author attributed about $120,000 in model spend and, by using his share of pull requests as a proxy, roughly three weeks of his own effort. That was not elapsed project time or total engineering effort: the work ran from May to August 2026 and other engineers contributed design, interoperation, packaging, build work and review.

GitHub also reported dozens of known regressions. Its useful rules were to keep end-to-end tests independent of the agent, prevent the agent from weakening the test oracle, translate before redesigning, and concentrate unsafe code at reviewable external boundaries. The source was TypeScript, already a memory-managed language, so the project does not measure the security return of replacing C or C++. It does show that AI can make a large port affordable while leaving compatibility, architecture and evidence as human responsibilities.

What a credible roadmap looks like

A vendor does not need to promise a clean-sheet rewrite of the entire product. NSA and CISA explicitly describe incremental adoption of memory-safe languages, including interoperability with existing code. NIST’s Secure Software Development Framework likewise gives organisations room to prioritise work by risk, cost, feasibility and applicability.

For each exposed component, a credible roadmap should publish:

The same roadmap should identify the components that will remain in C or C++ and explain why. Those components need stronger isolation, sustained testing and a smaller interface. “Too hard to rewrite” is a statement about scheduling, not a security boundary.

NetScaler 15.1 is a useful first step. Changing the appliance base is difficult and necessary work, and several individual vulnerability paths are fixed. The preview still contains PHP and Perl control-plane code and native packet and configuration services, while broad privilege and containment questions remain open.

A responsible migration wraps the old system more safely, retires dangerous glue, replaces the highest-risk native boundaries, and removes supported paths from the data plane to management services. Each stage should leave an attacker with less reach and a smaller reward.

The test of modernisation is what the next serious bug can reach, and whether the damage stops at the component boundary.

Sources and scope

The NetScaler observations apply to VPX 15.1 preview build 9.33 and the earlier 14.1 build 73.30 assessment. Later releases require fresh validation. Public programme status and metrics were checked on 5 October 2026; company-reported measurements are identified as such.

#Modernisation #Legacy Systems #Secure by Design #Memory Safety #AI-Assisted Development #NetScaler