2

I'm working at a small office,I have an application,it's generate a big text file with 14000 lines;

after each generate i must filter it and it's really boring;

I wanna write an application with java till I'll can handle it as soon as possible.

Please help me; I wrote an application with scanner (Of course with help :) ) but it's not good becase it was very slow;

For example it's my file :

SET CELL:NAME=CELL:0,CELLID=3;
SET LSCID:NAME=LSC:0,NETITYPE=MDCS,T32=5,EACT=FILTER-NOFILTER-MINR-FILTER-NOFILTER,ENSUP=GV2&NCR,MINCELL=6,MSV=PFR,OVLHR=9500,OTHR=80,BVLH=TRUE,CELLID=3,BTLH=TRUE,MSLH=TRUE,EIHO=DISABLED,ENCHO=ENABLED,NARD=NAP_STLP,AMH=ENABLED(3)-ENABLED(6)-ENABLED(9)

and I want this output (filter :)

CELLID :  3
ENSUP  :  GV2&NCR
ENCHO  :  ENABLED
MSLH   :  TRUE
------------------------
Count of CELLID : 2

which solution is the best and the fastest than the other ?

it's my source code :

public static void main(String[] args) throws FileNotFoundException {
        Scanner scanner = new Scanner(new File("i:\\1\\2.txt"));
        scanner.useDelimiter(";|,");
        Pattern words = Pattern.compile("(CELLID=|ENSUP=|ENCHO=)");

        while (scanner.hasNextLine()) {
          String key = scanner.findInLine(words);

          while (key != null) {
            String value = scanner.next();
            if (key.equals("CELLID=")) 
              System.out.print("CELLID:" + value+"\n");
             //continue with else ifs for other keys
              else if (key.equals("ENSUP="))
            System.out.print("ENSUP:" + value+"\n");

            else if (key.equals("ENCHO="))
            System.out.print("ENCHO:" + value+"\n");
            key = scanner.findInLine(words);
          }
          scanner.nextLine();
        }

}

Thank you very much indeed ...

skaffman
  • 398,947
  • 96
  • 818
  • 769
Freeman
  • 9,464
  • 7
  • 35
  • 58

2 Answers2

4

Since your code has performance issues, you first need to find bottle neck. You can profile it with profiler available with IDE you use.

However since your code is not high in computation but IO intensive, both in reading file and output using System.out.print, that is where I would suggest you to improve on for improving on file IO.

.

Replace this line of code

Scanner scanner = new Scanner(new File("i:\\1\\2.txt"));

.

With this lines of code

File file = new File("i:\\1\\2.txt");
BufferedReader br = new BufferedReader( new FileReader(file)  );
Scanner scanner = new Scanner(br);

Let us know if this helps.

.

Since previous solution did not helped much, I made few more changes to improve your code. You may have to correct errors in parsing if any. I was able to display output of parsing 392832 lines in approx 5 seconds. Original solution takes more than 50 seconds.

Chages are as below:

  1. Use of StringTokenizer instead of Scanner
  2. Use of BufferedReader for reading file
  3. Use of StringBuilder to buffer output

.

public class FileParse {

    private static final int FLUSH_LIMIT = 1024 * 1024;
    private static StringBuilder outputBuffer = new StringBuilder(
            FLUSH_LIMIT + 1024);
    private static final long countCellId;

    public static void main(String[] args) throws IOException {
        long start = System.currentTimeMillis();
        String fileName = "i:\\1\\2.txt";
        File file = new File(fileName);
        BufferedReader br = new BufferedReader(new FileReader(file));
        String line;
        while ((line = br.readLine()) != null) {
            StringTokenizer st = new StringTokenizer(line, ";|, ");
            while (st.hasMoreTokens()) {
                String token = st.nextToken();
                processToken(token);
            }
        }
        flushOutputBuffer();
        System.out.println("----------------------------");
        System.out.println("CELLID Count: " + countCellId);
        long end = System.currentTimeMillis();
        System.out.println("Time: " + (end - start));
    }

    private static void processToken(String token) {
        if (token.startsWith("CELLID=")) {
            String value = getTokenValue(token);
            outputBuffer.append("CELLID:").append(value).append("\n");
            countCellId++;
        } else if (token.startsWith("ENSUP=")) {
            String value = getTokenValue(token);
            outputBuffer.append("ENSUP:").append(value).append("\n");
        } else if (token.startsWith("ENCHO=")) {
            String value = getTokenValue(token);
            outputBuffer.append("ENCHO:").append(value).append("\n");
        }
        if (outputBuffer.length() > FLUSH_LIMIT) {
            flushOutputBuffer();
        }
    }

    private static String getTokenValue(String token) {
        int start = token.indexOf('=') + 1;
        int end = token.length();
        String value = token.substring(start, end);
        return value;
    }

    private static void flushOutputBuffer() {
        System.out.print(outputBuffer);
        outputBuffer = new StringBuilder(FLUSH_LIMIT + 1024);
    }

}

.

Update on ENSUP and MSLH:

To me it looks like you have switched ENSUP and MSLH in if statement as below. Hence you see "MSLH" value for "ENSUP" and vice a versa.

} else if (token.startsWith("MSLH=")) {
    String value = getTokenValue(token);
    outputBuffer.append("ENSUP:").append(value).append("\n");
} else if (token.startsWith("ENSUP=")) {
    String value = getTokenValue(token);
    outputBuffer.append("MSLH:").append(value).append("\n");
}
Gladwin Burboz
  • 3,519
  • 16
  • 15
  • Dear Gladwin Burboz Thanks for your reply but it's still very slow. – Freeman Jan 18 '10 at 06:13
  • Mike, let me know if above solution is faster. – Gladwin Burboz Jan 19 '10 at 00:22
  • Issue with this one is that if outputBuffer gets too big, it can cause OutOfMemoryError. Solution would be to flush output from time to time. Performance could further be enhanced by flushing output in seperate thread. I will update above solution further if I get some time. – Gladwin Burboz Jan 19 '10 at 00:27
  • Updated code to flush outputBuffer each time it exceeds FLUSH_LIMIT. – Gladwin Burboz Jan 19 '10 at 02:47
  • Dear Gladwin good job; and thank you very much indeed. it so faster and it has the best performance... – Freeman Jan 19 '10 at 07:43
  • I am glad that it worked fine... I am not sure what kind of count you want but I updated the code to print total count of each occurence of CELLID. Please see variable "countCellId" added to code. Hope that is what you need. – Gladwin Burboz Jan 19 '10 at 16:37
  • Thanks a lot ; Have a nice life :) – Freeman Jan 19 '10 at 20:06
  • Thanks to all those who voted for the answer. I am new to this forum and more votes/reputation helps. – Gladwin Burboz Jan 20 '10 at 15:46
3

Simple text filtering is probably easier to write in Perl (my choice because I've been using it for years) or Python (what I recommend to new people because it's a more modern language).

Paul Tomblin
  • 179,021
  • 58
  • 319
  • 408