Recursive BBCode Parsing

Question

I'm trying to parse BBCode in my script. Now, it works seamelessly, until I try to indent BBCode that's more than just bold or underline - such as spoiler, url, font size, etc. - then it screws up. Here's my code:

function parse_bbcode($text) {
    global $db;
    $oldtext = $text;
    $bbcodes = $db->select('*', 'bbcodes');
    foreach ($bbcodes as $bbcode) {
        switch ($bbcode->type) {
            case 'simple': {
                $find = '{content}';
                $replace = '${1}';
                $text = preg_replace(
                    '/\['.$bbcode->tag.'\](.+)\[\/'.$bbcode->tag.'\]/i',
                    str_replace($find, $replace, $bbcode->html),
                    $text);
                    break;
            }
            case 'property':
            case 'options': {
                $find = array ( '{property}', '{content}' );
                $replace = array ( '${1}', '${2}' );
                $text = preg_replace(
                    '/\['.$bbcode->tag.'\=(.[^\"]*)\](.+)\[\/'.$bbcode->tag.'\]/i',
                    str_replace($find, $replace, $bbcode->html),
                    $text);
                    break;
            }
        }
    }
    return $text;
}

Now my guess is that the RegEx doesn't like the recursiveness in the pattern. How can I improve it? A sample $bbcode object is as such:

stdClass::__set_state(array(
   'id' => '2',
   'name' => 'Italic',
   'type' => 'simple',
   'tag' => 'i',
   'button_image' => NULL,
   'button_text' => '<i>I</i>',
   'options' => '',
   'prompt' => NULL,
   'html' => '<i>{content}</i>',
   'order' => '1',
))
stdClass::__set_state(array(
   'id' => '3',
   'name' => 'URL',
   'type' => 'property',
   'tag' => 'url',
   'button_image' => NULL,
   'button_text' => 'http://',
   'options' => '',
   'prompt' => 'URL address',
   'html' => '<a href="{property}">{content}</a>',
   'order' => '4',
))

[PHP has a BBCode parser, so no reason to reinvent the wheel](http://php.net/manual/en/book.bbcode.php) — Gordon, Jul 21 '11 at 08:25
No shit. I didn't know that. How about making this an answer so I can mark it — casraf, Jul 21 '11 at 08:27
I added a few links to Clement's answer below. Feel free to accept that instead :) — Gordon, Jul 21 '11 at 08:38
For anyone looking, I'd like to add the library I wrote that is currently the best of those I was able to find: https://github.com/thunderer/Shortcode — Tomasz Kowalczyk, Dec 29 '15 at 13:53

score 7 · Accepted Answer · edited Jul 21 '11 at 08:39

7

As gordon said in comments PHP has a BBCode parser, so no reason to reinvent the wheel.

The native parser is a PECL package though, so you will have to install it. If that's not an option (for instance due to shared hosting), there is also a PEAR package: http://pear.php.net/package/HTML_BBCodeParser

In addition to those, you can take a look at forums using BB code source code, and either use their parser, or improve it. There is also several PHP implementations listed at http://www.bbcode.org/implementations.php

edited Jul 21 '11 at 08:39

Gordon

312,688
75
539
559

answered Jul 21 '11 at 08:28

Clement Herreman

10,274
4
35
57

The PECL package seems awesome, but I can't install it... The PEAR one is lame :/ – casraf Jul 21 '11 at 10:02

score 0 · Answer 2 · edited Dec 28 '14 at 04:54

0

Correctly parsing BBcode using regex is non-trival. The codes may be nested. CODE tags may contain BBCodes which must be ignored by the parser. Certain tags may not appear inside of other tags. etc. However, it can be done. I recently overhauled the BBCode parser for the FluxBB open source forum software. You may want to check it out in action:

New 2011 FluxBB Parser

Note that this new parser has not yet been incorporated into the FluxBB codebase.

edited Dec 28 '14 at 04:54

mario

144,265
20
237
291

answered Jul 21 '11 at 14:32

ridgerunner

33,777
5
57
69

Is this fully operational? And how easily can I add my own BBCode? – casraf Jul 21 '11 at 14:56
Yes it is operational (the link above is a test forum which uses this parser). This is not a generic parser - Its designed specifically to work with FluxBB. You'll need to read the [documentation](http://jmrware.com/articles/2011/fluxbb_jmr_dev/viewforum.php?id=1) and the [source code](https://github.com/jmrware/fluxbb_jmr_dev/blob/dev/include/parser.php). I assume you are a PHP programmer and are familiar with [regular expressions](http://jmrware.com/articles/2011/fluxbb_jmr_dev/viewtopic.php?id=3). As I said, parsing BBCode is non-trivial. – ridgerunner Jul 21 '11 at 23:27
Yeah I'm familiar with PHP and RegEx (see question), but BBCode is simple enough to parse -- until it gets recursive – casraf Jul 22 '11 at 10:36
@OhMrBigshot: The complexity of the problem is not limited to just recursion but in how the tags nest inside each other. e.g. Will you allow a `URL` tag inside an `IMG` tag? Or how about allowing a `QUOTE` tag inside an `I` tag? You need to set up a system of rules (which parallel HTML) that define what is allowed to nest inside each tag. Also, there needs to be rules for allowable parent tags, e.g. A `LI` (or `*`) tag may only appear inside a `LIST` tag. – ridgerunner Jul 22 '11 at 12:28

Recursive BBCode Parsing

2 Answers2

Linked