Until now there's been no way to judge whether a phe-estimator app is accurate. PKU Commons gives you the method, the data provenance, and a public benchmark, open to your review.
The Skill encodes the method a trained PKU parent uses, made explicit and reviewable.
This is the part built for you: a reproducible standard, verified independently of vendor claims.
Reproducible ground truth. Each benchmark case pairs a food label with a phe value computed from documented public sources, USDA FoodData Central and Open Food Facts, with the source id recorded. You (or anyone) can re-derive the "true" value independently.
Transparent metrics. Every estimator is scored the same way: mean absolute error, per-case error, and the share of estimates falling within a clinically-relevant tolerance band.
Versioned & audited. Methodology rules and food-list rows carry a citation to authority, a version, and a reviewer of record. Your challenges become tracked issues with documented resolutions, recorded in an audit trail anyone can inspect.
Your sign-off has standing. Clinician / RD review is a first-class authority layer: a named professional endorsement is a citable source in the record.
Phebe is a phe-estimator and logger. Try it on labels you know and compare against your own judgment.
Phebe is a reference implementation of the PKU Commons standard a working example of the open method and benchmark, not the standard itself. Any app can implement the standard.
What must a phe-estimator get right to be useful in your clinic or study? Where would an estimate be dangerous if wrong? Your input directly shapes the tolerance bands and the method.