coreml : fix --quantize crash for mlprogram format; fix --optimize-ane label (#3868)
commit
8b92060 switched ct.convert() to mlprogram, but did not update
the --quantize path. quantize_weights() from
neural_network.quantization_utils only works with the legacy
neuralnetwork format. Running with --quantize crashed with:
Exception: MLModel of type mlProgram cannot be loaded just from the
model spec object. It also needs the path to the weights file.
Fix: pass compute_precision=ct.precision.FLOAT16 into ct.convert() when
--quantize is set. This matches the original intent of nbits=16 (F16
storage) without changing the quantization scheme or model accuracy.
Also fix the three boolean CLI flags (--encoder-only, --quantize,
--optimize-ane) to use a _str_to_bool helper so that both
--flag True
and
--flag False
parse correctly. The type=bool form accepted "False" as True because
bool("False") == True.
Remove the "currently broken" label from --optimize-ane: the ANE path
(WhisperANE with Conv2d attention and LayerNormANE) converts and loads
correctly with both PyTorch 2.x and coremltools 9.x.