It's a fairly large code base split across 3 repos.
The good thing was that it is fairly easy to verify: we have a working (but slow) version that uses Spark, with lots of existing unit tests.
We verified by using those unit tests as well as running our end-to-end process in the Spark and Pandas version and verifying the two databases were within the differential-privacy noise bands of each other.
It's a fairly large code base split across 3 repos.
The good thing was that it is fairly easy to verify: we have a working (but slow) version that uses Spark, with lots of existing unit tests.
We verified by using those unit tests as well as running our end-to-end process in the Spark and Pandas version and verifying the two databases were within the differential-privacy noise bands of each other.