Skip to content

description: "Data Formats (CSV/JSON/XML) — Always pass newline='' when opening files for csv — it prevents extra blank lines on some platforms."---

05 · Data Formats (CSV/JSON/XML)

🎥 Video walkthrough

Most real programs spend a lot of time reading and writing data in standard interchange formats. Python's standard library ships solid support for the three you'll meet most often: CSV, JSON, and XML.

CSV — reading

import csv

# sample.csv:
# name,age,city
# Ada,36,London
# Grace,85,New York

with open("sample.csv", newline="") as f:
    reader = csv.reader(f)
    header = next(reader)         # first row
    for row in reader:
        print(row)                # each row is a plain list of strings
# ['Ada', '36', 'London']
# ['Grace', '85', 'New York']

CSV — DictReader for named columns

with open("sample.csv", newline="") as f:
    reader = csv.DictReader(f)
    for row in reader:
        print(row["name"], row["age"])   # each row is an OrderedDict/dict
# Ada 36
# Grace 85

CSV — writing

rows = [
    {"name": "Ada", "age": 36, "city": "London"},
    {"name": "Grace", "age": 85, "city": "New York"},
]

with open("output.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "age", "city"])
    writer.writeheader()
    writer.writerows(rows)

Always pass newline="" when opening files for csv — it prevents extra blank lines on some platforms.

JSON — reading and writing

JSON maps very naturally onto Python's built-in types.

import json

data = {
    "name": "Ada Lovelace",
    "born": 1815,
    "contributions": ["Analytical Engine notes", "first algorithm"],
    "active": False,
}

# Python object -> JSON string
text = json.dumps(data, indent=2)
print(text)

# JSON string -> Python object
parsed = json.loads(text)
print(parsed["contributions"][0])   # Analytical Engine notes

JSON — files directly

with open("person.json", "w") as f:
    json.dump(data, f, indent=2)

with open("person.json") as f:
    loaded = json.load(f)

JSON type mapping

JSON Python
object dict
array list
string str
number int or float
true/false True/False
null None

Handling malformed JSON

try:
    json.loads("{not valid json}")
except json.JSONDecodeError as e:
    print(f"invalid JSON at line {e.lineno}, column {e.colno}: {e.msg}")

XML — parsing with ElementTree

import xml.etree.ElementTree as ET

xml_text = """
<library>
    <book id="1">
        <title>Structure and Interpretation</title>
        <author>Abelson</author>
    </book>
    <book id="2">
        <title>Fluent Python</title>
        <author>Ramalho</author>
    </book>
</library>
"""

root = ET.fromstring(xml_text)

for book in root.findall("book"):
    title = book.find("title").text
    author = book.find("author").text
    print(f"{title} by {author} (id={book.get('id')})")
# Structure and Interpretation by Abelson (id=1)
# Fluent Python by Ramalho (id=2)

XML — building and writing

library = ET.Element("library")
book = ET.SubElement(library, "book", id="3")
ET.SubElement(book, "title").text = "Automate the Boring Stuff"
ET.SubElement(book, "author").text = "Sweigart"

tree = ET.ElementTree(library)
tree.write("library.xml", encoding="utf-8", xml_declaration=True)

Choosing a format

Format Good for Watch out for
CSV tabular data, spreadsheets no nested structure, everything is a string
JSON nested data, config, web APIs no comments, no native dates
XML documents, legacy enterprise systems, attributes + text more verbose, easy to get parsing wrong

How It Actually Works

Each parser is a small state machine turning a byte/character stream into Python objects — they differ in what grammar they implement:

  • csv is a C-level character-by-character state machine implementing RFC 4180. It tracks states like start-of-field, in-quoted-field, in-unquoted-field, quote-in-quoted-field. That's why you can't just line.split(","): a quoted field may contain commas, embedded newlines, or doubled quotes ("""), and only the state machine handles them correctly. DictReader runs the same machine, then zips the first row (header) against each subsequent row.
  • json tokenizes the text ({, }, [, ], :, ,, strings, numbers, true/false/null) and recursively builds objects: a { starts a dict, keys/values fill it until }. CPython ships a C accelerator (_json) that does this in one pass; there's a pure-Python fallback with identical behavior. json.dumps is the reverse recursion, emitting text for each object and raising TypeError on anything it doesn't know how to represent (e.g. a datetime). The JSONDecodeError carries lineno/colno because the scanner tracks its position as it goes.
  • xml.etree.ElementTree drives the C expat library, a streaming (SAX-style) parser that fires callbacks — start tag, text, end tag — as it reads. ET.fromstring uses a TreeBuilder that responds to those callbacks by constructing Element nodes and nesting them, giving you the finished tree. This streaming core is why iterparse can process gigabyte-scale XML without loading it all into memory.

Exercise

Given a CSV file of products (name,price,quantity), write a script that: reads it with csv.DictReader, converts price to float and quantity to int, computes each product's total = price * quantity, and writes the result as a JSON file — a list of objects with name and total — sorted by total descending.