Targeted data poisoning of neural radiance fields
Problem. A NeRF learns a 3D scene from a set of photos. If an attacker can edit some of those photos, can they remove one object from the scene, and what does that do to everything else? To measure it I built a Blender pipeline for a procedural scene and camera rig, a scripted compositor that erases the target from selected views, retraining, and evaluation of the target region and the rest of the frame separately, on a held-out view set that was frozen and checksummed before the first poisoned run.
Key decision. The methodology was locked before any results existed, and every non-trivial choice since is logged with the alternatives I rejected. That discipline paid off before the main sweep: it caught four silent config defects, including every poisoned condition quietly training on clean data, before any compute was spent on them.
Result. The proof of concept passed: target‑region PSNR fell from 23.6 to 19.4 to 11.6 dB as the poisoned share went from 0 to 20 to 50% (Figure 1). Next is an 8-condition, 3-seed sweep at 200,000 iterations per run.