A comprehension building 2M items is slower than the equivalent loop. Expected?
Not usually, so look at what is inside. A comprehension that calls a method per item pays the attribute lookup every time; hoisting it out (append = out.append) or using a generator when you do not need the list is what actually moves the number.
A 300 MB JSON array does not fit comfortably in memory. Options?
Stream it. Either a pull parser like ijson, or, if the structure is a flat array of objects, read in chunks and use json.JSONDecoder().raw_decode to peel one object at a time. Memory then depends on the largest single object, not on the file.
Is there a real reason to use dataclasses over dicts for internal data?
Typos become errors instead of silent None, and the field list is documentation that cannot drift. The cost is a small allocation overhead, which matters only in the hottest loops. With slots=True even that mostly disappears.