Verkracked Part 8 – VerkEye

I recently put together a local visual-recognition harness for the Verkada CB62 cameras I have been researching.

I wanted to be able to see what the camera sees and recognizes without keeping the physical device online, talking to the cloud, or tied up on my bench. The goal was basically the same thing that pushed me to build BirdEye: give the device an image, a video, or a live webcam feed and watch its own recognition pipeline work locally.

I assumed the recovered model would eventually turn into a fairly ordinary conversion project. It did not.

The model is not TensorFlow, TFLite, ONNX, or a normal collection of weights. It is a proprietary, ten-split package compiled for Ambarella’s CV22 vision accelerator, surrounded by AArch64 userspace libraries, a Cavalry kernel interface, firmware, microcode, and device-specific production configuration.

Those binaries were written to run inside the camera. macOS does not have a Cavalry accelerator. A normal Linux workstation does not have one either. QEMU could execute the recovered ARM userspace and get all the way to the real CAVALRY_RUN_DAGS request, but there was no CV22 hardware behind that ioctl to perform the actual inference.

So this was not a matter of recompiling a library. It became a compatibility-runtime project.

What it actually took

I first had to map the package and its boundary ABI: ten compiled split records, 39 tensor descriptors, 16 connections between splits, and six terminal detector outputs. Then I had to recover the camera’s real preprocessing contract, the detector geometry, the class mapping, the thresholds, the filtering, and the non-maximum suppression behavior.

  1. Recover: the CV22 model package and vendor runtime.
  2. Map: ten compiled splits, 39 tensor descriptors, 16 split edges, and six terminal outputs.
  3. Oracle: preserve Ambarella ADES/Cavalry execution as an independent source of raw outputs.
  4. Rebuild: implement compatible MLX and OpenVINO execution paths for macOS and Linux.
  5. Prove: compare raw tensors and production detections byte-for-byte.

The recovered model graph ultimately required 55 fast-convolution and mixed operators. The awkward part was not just finding familiar neural-network shapes. It was reproducing CV22-specific behavior closely enough that the results were literally the same: dynamic SPPF max pooling, signed 2×2 transpose-fastconv accumulator saturation, the pre-offset accumulator clamp, and multiple recovered sigmoid mappings all mattered.

I kept the Ambarella ADES/Cavalry execution path as an independent oracle and compared the replacement runtime at every terminal head. When a natural traffic-camera frame exposed 65 drifting score elements, that was not waved away as “close enough.” The mismatch stayed a failure until the underlying arithmetic was corrected.

  • 29 registered synthetic and natural test cases.
  • 174 terminal tensors compared byte-for-byte.
  • 0 differing bytes across both accelerated backends.
  • 32.236698 FPS on the tested Apple-silicon MLX path.

The final corpus covers constants, boundary values, seeded random input, structured patterns, generated media, 18 public traffic-camera images, and decoded-video frames. Both the macOS MLX backend and the Linux OpenVINO backend match all 174 terminal tensors with zero differing bytes, then match the assembled 13,566 × 8 prediction matrix and production postprocessor output.

That is the important distinction: VerkEye does not swap in a convenient YOLO model, retrain replacement weights, or invent an ONNX graph that looks about right. It executes the exact recovered CB62 artifact through a reconstructed compatibility layer and fails closed if the model, tensor ABI, geometry, or evidence does not match.

What it does now

VerkEye can inspect the recovered package, run local still images and recorded video, process frame directories, and use a webcam in a live native viewer. The viewer draws the recovered person, vehicle, and animal classes with confidence, frame counters, measured FPS, pause/resume controls, and optional annotated recording.

Demo 1: a refreshed Singapore traffic-camera source using the explicit small-object profile. The model bytes and confidence thresholds remain unchanged.

Demo 2: a bundled offline street scene using the unchanged production profile, with two accepted person detections.

On the tested Apple-silicon Mac, the complete model-boundary path—preprocessing, recovered graph, prediction assembly, and production postprocessor—measured 32.236698 FPS. That makes the local viewer genuinely usable in real time.

The tested Linux path is exact too, but the Intel UHD 630 OpenVINO host only reached 3.051105 FPS. That is useful for offline analysis, but I am not pretending it is a good live-video experience on that hardware. Faster Linux hardware is still an open performance target; correctness is already gated the same way on both platforms.

Release boundary

The distributable project includes the VerkEye code, the derived compatibility runtime, the GUI, the demos, installation tooling, and the verification machinery. It intentionally does not include the recovered Verkada model or the Ambarella/Verkada executables, libraries, kernel module, firmware, microcode, or production configuration.

To use the exact CB62 path, you bring yolov6n_hor.bin from a CB62 you own. VerkEye hash-checks it before execution. This keeps the interesting part—the independently built runtime and research tooling—usable without redistributing the proprietary artifact that came from the camera.

The public project is available at github.com/GainSec/VerkEye.

It was definitely a much larger lift than I expected when I started. What looked like “port the model” became a multi-layer reconstruction of a proprietary accelerator pipeline, its arithmetic edge cases, and the application around it. But now I can point a normal webcam or video file at the exact CB62 model on a Mac, see the boxes and labels in real time, and prove that the six raw detector heads are the same bytes the vendor runtime produced.

View other parts of the Verkracked Research Project Below:

Part 0 – Verkracked – Security Research on Verkada Anti-Crime Devices – Link
Part 4&5 – Verkracked – Local Cloud and Sub-GHz Frameworks for Verkada Alarm Hubs – LinkWe will see what comes from it.

END TRANSMISSION

Leave a Reply