How can I get all the data from separated large files in R Revolution Enterprise?

Question

I'm using RevoR entreprise to handle impoting large data files. The example given in the documentation states that 10 files (1000000 rows each) will be imported as dataset using an rxImport loop like this :

setwd("C:/Users/Fsociety/Bigdatasamples")
Data.Directory <- "C:/Users/Fsociety/Bigdatasamples"
Data.File <- file.path(Data.Directory,"mortDefault")
mortXdfFileName <- "mortDefault.xdf"

append <- "none"
for(i in 2000:2009){
importFile <- paste(Data.File,i,".csv",sep="")
mortxdf <- rxImport(importFile, mortXdfFileName, append = append, overwrite = TRUE, maxRowsByCols = NULL)
append <- "rows"    
}
mortxdfData <- RxXdfData(mortXdfFileName)
knime.out <- rxXdfToDataFrame(mortxdfData)

The issue here is that I only get 500000 rows in the dataset due to the maxRowsByCols argument the default is 1e+06 i changed it to a higher value and then to NULL but it still truncates the data from the file.

score 1 · Answer 1 · answered Mar 17 '16 at 19:56

Since you are importing to an XDF the maxRowsByCols doesn't matter. Also, on the last line you read into a data.frame, this sort of defeats the purpose of using an XDF in the first place.

This code does work for me on this data http://packages.revolutionanalytics.com/datasets/mortDefault.zip, which is what I assume you are using.

The 500K rows is due to the rowsPerRead argument, but that just determines block size. All of the data is read in, just in 500k increments, but can be changed to match your needs.

setwd("C:/Users/Fsociety/Bigdatasamples")
Data.Directory <- "C:/Users/Fsociety/Bigdatasamples"
Data.File <- file.path(Data.Directory, "mortDefault")
mortXdfFileName <- "mortDefault.xdf"

append <- "none"
overwrite <- TRUE
for(i in 2000:2009){
  importFile <- paste(Data.File, i, ".csv", sep="")
  rxImport(importFile, mortXdfFileName, append=append, overwrite = TRUE)
  append <- "rows"
  overwrite <- FALSE
}

rxGetInfo(mortxdfData, getBlockSizes = TRUE)

# File name: C:\Users\dnorton\OneDrive\R\MarchMadness2016\mortDefault.xdf 
# Number of observations: 1e+07 
# Number of variables: 6 
# Number of blocks: 20 
# Rows per block (first 10): 5e+05 5e+05 5e+05 5e+05 5e+05 5e+05 5e+05 5e+05 5e+05 5e+05
# Compression type: zlib

Thanks i'll try this out. But i'm using the xdf file as a reduction to the huge data i'm going to import. and the data.frame parse is for Knime nodes (they can't handle rxXDFData class only data.frame) — Ismail Akrim, Mar 18 '16 at 15:32

score 1 · Accepted Answer · answered Mar 18 '16 at 15:45

1

Fixed, the issue was that the RxXdfData() has a maxrowbycols limitation,changing it to NULL will convert the whole rxXdfData into a data.frame object for Knime.

answered Mar 18 '16 at 15:45

Ismail Akrim

51
7

How can I get all the data from separated large files in R Revolution Enterprise?

2 Answers2