Pandas Memory Error

May 03, 2024 Post a Comment

I have a csv file with ~50,000 rows and 300 columns. Performing the following operation is causing a memory error in Pandas (python): merged_df.stack(0).reset_index(1) The data fr

Solution 1:

So it takes on my 64-bit linux (32GB) memory, a little less than 2GB.

In [5]: def f():
       df = DataFrame(np.random.randn(50000,300))
       df.stack().reset_index(1)


In [6]: %memit f()
maximum of 1: 1791.054688 MB per loop

Since you didn't specify. This won't work on 32-bit at all (as you can't usually allocate a 2GB contiguous block), but should work if you have reasonable swap / memory.

Solution 2:

As an alternative approach you can use the library "dask" e.g:

Baca Juga

Python Line-by-line Memory Profiler?
Function Should Clean Data To Half The Size, Instead It Enlarges It By An Order Of Magnitude
Add New Row Based On An If Condition Via Python

# Dataframes implement the Pandas API
import dask.dataframe as dd`<br>
df = dd.read_csv('s3://.../2018-*-*.csv')

Python Guru

Pandas Memory Error

Solution 1:

Solution 2:

Post a Comment for "Pandas Memory Error"