What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For a large JSON dataset, use newline-delimited JSON (JSON Lines), read it with pd.read_json(..., lines=True, chunksize=...), and process each chunk without retaining the whole file. Reduce columns and choose compact dtypes before doing any work. Pandas remains an in-memory tool, so global joins, sorts, and groupby operations may still exceed RAM; those workloads need an out-of-core or distributed engine.
What pandas can and cannot do with data larger than RAM
Pandas stores DataFrames in memory. A file that appears to fit in available RAM can still fail when parsing creates temporary objects, type conversions make copies, or an operation builds a second DataFrame. The practical goal is to lower peak memory, not merely the size of the source file.
- Load only the fields needed for the analysis.
- Declare or convert dtypes deliberately, especially for identifiers and low-cardinality text.
- Process independent or associative work one chunk at a time.
- Avoid concatenating every chunk unless the combined result is known to fit in memory.
The pandas 3.0.6 scaling guide illustrates the potential, not a universal benchmark: selecting four Parquet columns used about one-tenth of the memory in its example. In another example, converting a low-cardinality text field to category and downcasting numeric columns produced a displayed memory ratio of 0.42; the guide says the in-memory footprint fell to one-fifth of its original size. Actual savings depend on cardinality, nulls, values, and whether later operations create copies.
Apply memory controls before parsing
Read fewer fields
For JSON, select fields during normalization when possible. For CSV or other staging files, usecols prevents unneeded columns from becoming part of the DataFrame:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
- FULL HD IPS DISPLAY - Enjoy vibrant, crystal-clear images with 178-degree wide-viewing angles
- AMD RYZEN 3 30 PROCESSOR - Everyday performance you can count on; Multitask, stream, game casually, and edit photos smoothly with responsive power and vibrant HDR visuals
- ENJOY UP TO 14 HOURS AND 15 MINUTES OF BATTERY LIFE - HP Fast Charge restores battery from 0 to 50% in approximately 45 minutes
- AMD RADEON 610M GRAPHICS - Experience smooth entertainment; Built for streaming and multitasking, enjoy realistic visuals and efficient performance for work and play
- STORAGE AND MEMORY - 512 GB PCIe NVMe M.2 SSD offers fast speed and efficient storage; and 8 GB LPDDR5 RAM memory boosts performance with higher bandwidth
import pandas as pd
sales = pd.read_csv(
'sales.csv',
usecols=['account_id', 'region', 'amount'],
dtype={'account_id': 'string', 'region': 'category', 'amount': 'float32'}
)
low_memory=True changes CSV parser internals; it does not make the final DataFrame out-of-core. Without chunksize or iterator, pandas still returns one complete DataFrame.
Protect identifiers and choose numeric widths
Keep ZIP codes, account numbers, and other identifiers as strings when leading zeros or nonnumeric characters matter. Downcast only after checking the valid range and missing-value behavior. Nullable pandas or PyArrow dtypes can preserve missing values without forcing an inappropriate Python object column.
Stream newline-delimited JSON with chunks
JSON Lines (often called JSONL) stores one complete JSON object per line. It is the practical JSON format for pandas iteration:
import pandas as pd
reader = pd.read_json('events.jsonl', lines=True, chunksize=100_000)
for chunk in reader:
# Work on this chunk, then release it before reading the next one.
print(len(chunk))
With lines=True and chunksize, pandas returns a JsonReader iterator. The chunk size is a working parameter, not a guaranteed safe number: reduce it if parsing or downstream operations peak too high, and increase it only after measuring memory and throughput on representative data.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
- Intel Celeron N4120: 4 Cores & Threads, 1.1GHz Base Clock, Up to 2.6GHz Boost Clock, 4MB Cache, Intel UHD Graphics 600. The perfect combination of performance, power consumption, and value helps your device handle multitasking smoothly and reliably with four processing cores to divide up the work.
- 14" HD Display: 14.0-inch diagonal, HD (1366 x 768), micro-edge, anti-glare. See your digital world in a whole new way. Enjoy movies and photos with the great image quality and high-definition detail of 1 million pixels.
- Memory & Storage: 4 GB LPDDR4x & 64 GB eMMC Storage. Adequate high-bandwidth RAM to smoothly run multiple applications and browser tabs all at once. An embedded multimedia card provides reliable flash-based storage.
- Ports:2 x USB 3.0 Type-A,1 x USB 3.0 Type-C,1 x HDMI,1 x Headphone Jack
- Chrome OS: Chromebook is a computer for the way the modern world works, with thousands of apps. Enjoy the seamless simplicity that comes with Google Chrome and Android apps, all integrated into one laptop. It’s fast, simple, and secure.
Aggregate without rebuilding the full table
Chunking is strongest when each chunk can be processed independently and partial results can be combined with an associative operation. This example counts event types while parsing timestamps and retaining missing event types:
import pandas as pd
counts = None
for chunk in pd.read_json('events.jsonl', lines=True, chunksize=100_000):
chunk['event_time'] = pd.to_datetime(
chunk['event_time'], errors='coerce', utc=True
)
part = chunk.groupby('event_type', dropna=False).size()
counts = part if counts is None else counts.add(part, fill_value=0)
counts = counts.astype('int64').rename('count').reset_index()
Define the aggregation’s missing-value policy and verify that combining partial results gives the same answer as a single pass. Keep only the aggregate, not every processed chunk.
When the source is a regular JSON array
A conventional JSON document commonly contains one large array surrounded by a single pair of brackets. The line-oriented iterator pattern does not apply directly to that layout; loading and decoding the document can require the whole object in memory. If you control the export, produce JSON Lines. Otherwise, convert the array with a streaming-aware preprocessing step or use a parser designed for incremental JSON before handing manageable pieces to pandas.
Flatten nested JSON deliberately
pd.read_json handles file-level JSON parsing and iteration. pd.json_normalize turns nested records into tabular columns. Use the latter when objects contain nested dictionaries or arrays:
Rank #3
- Stunning 15.6" FHD IPS Display: Experience crisp 1920x1080 resolution on this 15.6 inch laptop with an IPS panel that delivers wide viewing angles and vivid colors. The narrow-bezel design maximizes screen real estate for comfortable viewing on this Win 11 laptop, whether you're studying or working.
- Celeron J4105 Processor & 256GB SSD: Powered by a reliable Celeron J4105 processor paired with 12GB DDR4 memory and a fast 256GB M.2 SSD. This laptop computer supports SSD expansion up to 2TB and TF card expansion up to 1TB, so your storage grows with your needs. Delivers smooth multitasking for daily productivity.
- AI-Powered Win 11 Laptop: Built-in AI features enhance your productivity with smart assistance for writing, summarizing, and task management. Pre-installed with Win 11 and includes Office 365 subscription. This student laptop is backed by 1-year warranty and 24/7 customer support.
- All-Day 7000mAh Battery & 180° Hinge: The high-capacity 7000mAh battery keeps this laptop powered through long classes or meetings. The 180-degree lay-flat hinge lets you share your screen effortlessly during presentations. This durable laptop computer adapts to your dynamic workflow.
- Versatile Connectivity Hub: Equipped with USB 3.2, Type-C, Mini HDMI, and 3.5mm audio jack to connect all your peripherals. Stay online anywhere with high-speed 5G WiFi and Bluetooth 4.2. This college laptop keeps you connected at home, in the library, or on the go.
import json
import pandas as pd
with open('orders.json', encoding='utf-8') as file:
payload = json.load(file)
orders = pd.json_normalize(
payload,
record_path='orders',
meta=['customer_id', 'created_at'],
sep='.'
)
record_pathidentifies the list whose members become rows.metacopies parent-level fields onto those rows.sep='.'creates predictable names such asshipping.city.- Missing keys become missing values; validate required fields after normalization.
Account for list expansion
Flattening a nested list changes the table’s grain. One order with five line items becomes five rows; exploding a second list can create a multiplicative combination. State whether a row represents an order, an order item, or another entity before joining the result back to a parent table:
items = pd.json_normalize(payload, record_path='orders', sep='.')
items = items.explode('tags', ignore_index=True)
Do not silently aggregate away the extra rows. If arrays represent independent entities, normalize them into separate tables keyed by the parent identifier.
Choose the JSON orientation that matches the producer
| Orientation | Layout | Use and limitation |
|---|---|---|
records |
List of row objects | Natural for records and JSON Lines; index labels are not preserved. |
split |
Separate columns, index, and data arrays | Preserves index and column structure explicitly. |
index |
Object keyed by index | Useful when row labels are the outer keys. |
columns |
Object keyed by column | Column-oriented representation. |
values |
Nested value arrays only | Compact, but carries no labels. |
table |
Schema plus data section | Suitable when the producer supplies a table schema and data. |
Pass the matching orient when reading or writing a non-default document. Do not infer semantic meaning from automatic date conversion: parse units, time zones, and invalid values explicitly when the data contract requires it.
Use PyArrow where its support fits
Pandas can expose nullable columns backed by Apache Arrow with dtype_backend='pyarrow':
Rank #4
- Efficient Performance for Everyday Computing: Powered by Intel N150 processor with up to 3.6 GHz Intel Turbo Boost Technology, 6 MB L3 cache, 4 cores, and 4 threads, this HP laptop delivers responsive performance for web browsing, streaming, document editing, and multitasking. Paired with 4GB LPDDR5 RAM and 128GB UFS storage, it handles daily tasks smoothly. Includes 1-year Microsoft 365 Personal subscription for Word, Excel, PowerPoint, and cloud storage to maximize your productivity.
- 14-Inch HD Micro-Edge Display:Enjoy clear visuals on the 14-inch HD (1366 x 768) anti-glare screen with 250-nit brightness and 62.5% sRGB coverage. The micro-edge bezel delivers a 79% screen-to-body ratio in a compact design. An HP True Vision 720p HD camera with noise reduction and dual-array microphones supports clear video calls, remote work, and online learning.
- Modern Connectivity and Wireless Technology: Stay connected with Wi-Fi 6 (2x2) for faster wireless speeds and Bluetooth 5.4 for seamless pairing with accessories. Versatile port selection includes 1 USB Type-C 10Gbps with DisplayPort 1.2 for external displays, 2 USB Type-A 5Gbps ports for peripherals, 1 HDMI 1.4b port, 1 headphone/microphone combo jack, and 1 multi-format SD media card reader. Connect monitors, transfer files quickly, and expand your workspace with ease.
- All-Day Battery Life and Portable Design: Enjoy up to 11 hours of video playback, 7.5 hours of mixed usage, or 7.5 hours of wireless streaming on a single charge, perfect for students and professionals on the go. Weighing just 3.24 lb and measuring 12.76" x 8.86" x 0.71", this lightweight laptop fits easily in backpacks and bags. The stylish willow green top cover with matte finish and natural silver keyboard deck with vertical brushing pattern offer a modern, professional look.
- AI-Enhanced Productivity: Access Microsoft Copilot instantly with the dedicated Copilot key for faster assistance. AI Noise Reduction filters background sounds and improves voice clarity during calls. Dual speakers provide clear audio, while the full-size natural silver keyboard and HP Imagepad support comfortable typing and navigation.
import pandas as pd
events = pd.read_json(
'events.jsonl',
lines=True,
chunksize=100_000,
dtype_backend='pyarrow'
)
Arrow-backed columns can improve interoperability and sometimes reduce object-heavy memory use, but they do not remove pandas’ in-memory execution model. PyArrow is also an IO engine for supported pandas readers, yet engine coverage is reader- and option-specific. The pandas IO documentation notes that some features are unsupported in the PyArrow engine and that chunking behavior can differ by engine. Check the exact reader and pandas version before depending on an engine-specific option, and test dtypes after parsing.
Know when chunking is the right algorithm
| Workload | Chunking fit | Recommended approach |
|---|---|---|
| Per-record validation or file conversion | Strong | Process and write each chunk, keeping only counters and error records. |
| Value counts, sums, and other associative reductions | Strong | Compute a partial result and merge it with add, sums, or another proven associative combine. |
| Global groupby with manageable key state | Conditional | Aggregate partial groups, then combine; monitor the number of unique keys. |
| Joins requiring all keys | Weak | Use a keyed, out-of-core strategy or an engine that can coordinate partitions. |
| Global sorting | Weak | Use an external-sort or query engine rather than concatenating all chunks. |
| Algorithms requiring repeated passes | Weak | Choose a system designed for out-of-core or distributed execution. |
Pandas describes chunking as effective when an operation requires zero or minimal coordination between chunks. A chunked read alone cannot make a globally coordinated operation memory-safe.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A practical workflow for a large JSON analysis
- Confirm the format. Identify whether the input is JSON Lines, a regular array, or a nested document, and record the expected row grain.
- Inspect a small sample. Check key names, nulls, identifier formats, timestamp units, and nested-list lengths before selecting dtypes.
- Reduce the schema. Normalize only required fields; keep identifiers as strings and choose compact numeric or categorical types where valid.
- Set a conservative chunk size. Start small enough that parsing plus the planned operation fits comfortably in RAM, then measure.
- Process and release. Perform the per-chunk transformation, merge only the necessary partial result, and avoid a list of retained DataFrames.
- Validate totals. Compare row counts, null counts, key coverage, and representative values against the source or a smaller full-load sample.
- Escalate deliberately. If the required operation needs global coordination or repeated passes, move to an out-of-core, parallel, or distributed engine instead of endlessly tuning
chunksize.
Troubleshoot common failures
MemoryError or the process is killed
- Lower
chunksizeand remove unused fields before parsing. - Inspect
df.dtypesand convert object-heavy columns to appropriate string, categorical, nullable, or Arrow-backed types. - Look for accidental copies from chained transformations, concatenation, sorting, or joins.
- Write intermediate aggregates rather than retaining every chunk.
Numbers or identifiers are wrong
Leading zeros disappear when an identifier is parsed as an integer. Supply an explicit string dtype where supported and validate a sample after reading. Mixed numeric and text values should be handled as a documented schema decision, not left to inference.
Dates parse inconsistently
Date inference is a convenience, not proof that a timestamp has the intended timezone or unit. Use pd.to_datetime with explicit error handling and timezone treatment, then check invalid and ambiguous values.
Best Value
- Designed for mobility with a slim 0.71-inch profile and lightweight 3.24 lb chassis, making it easy to carry between home, office
Row counts increase unexpectedly
Inspect every nested array passed through record_path or explode. A list expansion legitimately creates multiple rows, but a second expansion or an incorrect metadata join can multiply them again. Validate the key and grain after each normalization step.
Chunking does not reduce memory
Confirm that the code is iterating over a JsonReader and not calling pd.concat on all chunks. Also check whether a later global operation recreates the entire dataset; the read phase may be bounded while the analysis phase is not.
When to leave pandas
Stay with pandas when the selected columns and required intermediate state fit in memory and the computation can be expressed as independent chunk work or a small associative reduction. Consider an out-of-core or distributed dataframe/query system when the job needs global joins, global sorting, repeated scans, or parallel execution across partitions. PyArrow-backed pandas columns improve type interoperability, but they are not by themselves a distributed execution engine.
Bottom line
For large JSON, prefer JSON Lines, read with lines=True and chunksize, normalize only the fields you need, and combine partial results instead of concatenating raw chunks. Treat dtypes, nested-list grain, timestamp semantics, and parser-engine support as correctness concerns as well as memory concerns. If the algorithm requires coordination across the entire dataset, use a system designed for that workload rather than forcing pandas to hold it all.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

