I like your code and have no problem with it!
There are a number of techniques to deal with this kind of parsing. I know how to implement several of them and I'm ok with them all.
Your example data format is unusual because it has more than one significant complicating factor.
Just for fun, I show an alternate coding that demo's some other techniques. I make no claim about "better". There is seldom a coding pattern that works "the best" in all situations. I used your regex'es as they looked fine to me. At the end of the day, all of the "states" have to be described and handled.
# Don't read in another line if we are still working
# on a START line. This is caused by the
# X START Y syntax in conjunction with the idea
# of END absent a START in this example file format.
# As a thought redefining the input separtator to
# be 'START' could possibly be productive if the format
# is not exactly like this?,
# This format has some of the nastiest things to deal
# with. They normally do not occur all at once!
my $line_in ='';
while ( $line_in =~ /START/ or $line_in =<DATA>)
$line_in = construct_record($line_in) if $line_in =~ /START/;
my $line = shift;
if ( (my $x) = $line =~ /START\s+(\w+)\s*$/)
push @record, $x;
while (defined ($line = <DATA>) and ($line !~ /(START|END)/) )
$line =~ s/^\s*|\s*$//g;
push @record, $line;
$line //= ''; #could be an EOF
if (my ($b4end) = $line =~ /^ (?: (.+) \s+ )? END $/x)
push @record, $b4end if $b4end;
return ''; # no continuation of this record
if ( my ($x,$y) = $line =~ /^ (?: (.+) \s+ )? START (?: \s+ (.+) )
+? $/x )
push @record, $x;
output_record(); # might be: "^START 77"?
return "START $y";
sub output_record # or process it somehow...
print "Record: @record\n" if (@record >1);
Record: a b
Record: c d e f g
Record: h i
Record: j k
i START j
Are you posting in the right place? Check out Where do I post X? to know for sure.
Posts may use any of the Perl Monks Approved HTML tags. Currently these include the following:
<code> <a> <b> <big>
<blockquote> <br /> <dd>
<dl> <dt> <em> <font>
<h1> <h2> <h3> <h4>
<h5> <h6> <hr /> <i>
<li> <nbsp> <ol> <p>
<small> <strike> <strong>
<sub> <sup> <table>
<td> <th> <tr> <tt>
Snippets of code should be wrapped in
<code> tags not
<pre> tags. In fact, <pre>
tags should generally be avoided. If they must
be used, extreme care should be
taken to ensure that their contents do not
have long lines (<70 chars), in order to prevent
horizontal scrolling (and possible janitor
Want more info? How to link or
or How to display code and escape characters
are good places to start.