I have a data frame (derived from a CSV file) with about 100M entries that looks like this: df1: var1 var2 0 1 2 1 2 1 2 1 {3,4,5} ...