Calibration exists because human QA drifts. Two analysts score the same call differently; the same analyst scores differently on a Friday afternoon. Regular calibration sessions keep a team scoring to the same standard and surface where the scorecard itself is ambiguous.
It is also a good test of an automated scorer: run it against a call the team has already calibrated on and see where it lands.
We do not ship a calibration workflow. Automated scoring removes the between-reviewer drift that calibration exists to correct, but if you run formal calibration sessions today, you would keep running them outside the product.