perception, hardware
Some of the most important information on a road has almost no shape.
You can often recognize someone you know from a grainy photo — not from a tape measure, but from posture, gait, and context. No single feature is precise. The recognition is a pattern.
Cameras capture that kind of information. A ranging laser, by design, does not. A point cloud describes shape and distance well and is quiet about almost everything else.
No single feature is precise.
Think about what a driver has to read. The color of a traffic light. Whether brake lamps are on. Text on a temporary sign. A person in a vest waving traffic through a red signal. Whether a pedestrian is looking at the roadway or at a phone. A turn indicator blinking. Skid marks. The difference between a shadow and a hole. Most of those cues have no distinctive geometry. A lit brake lamp and an unlit one are the same shape.
That is why a camera-first design is not only a cost story. Roads were built for eyes. A lot of the official meaning of a street is painted, lit, written, or gestured.
The honest trade-off remains: cameras infer distance instead of measuring it, and they depend on available light. Those are real weaknesses. The strategic question is which weakness is easier to engineer around.