I like your code and have no problem with it!
There are a number of techniques to deal with this kind of parsing. I know how to implement several of them and I'm ok with them all.
Your example data format is unusual because it has more than one significant complicating factor.
Just for fun, I show an alternate coding that demo's some other techniques. I make no claim about "better". There is seldom a coding pattern that works "the best" in all situations. I used your regex'es as they looked fine to me. At the end of the day, all of the "states" have to be described and handled.
# Don't read in another line if we are still working
# on a START line. This is caused by the
# X START Y syntax in conjunction with the idea
# of END absent a START in this example file format.
# As a thought redefining the input separtator to
# be 'START' could possibly be productive if the format
# is not exactly like this?,
# This format has some of the nastiest things to deal
# with. They normally do not occur all at once!
my $line_in ='';
while ( $line_in =~ /START/ or $line_in =<DATA>)
$line_in = construct_record($line_in) if $line_in =~ /START/;
my $line = shift;
if ( (my $x) = $line =~ /START\s+(\w+)\s*$/)
push @record, $x;
while (defined ($line = <DATA>) and ($line !~ /(START|END)/) )
$line =~ s/^\s*|\s*$//g;
push @record, $line;
$line //= ''; #could be an EOF
if (my ($b4end) = $line =~ /^ (?: (.+) \s+ )? END $/x)
push @record, $b4end if $b4end;
return ''; # no continuation of this record
if ( my ($x,$y) = $line =~ /^ (?: (.+) \s+ )? START (?: \s+ (.+) )
+? $/x )
push @record, $x;
output_record(); # might be: "^START 77"?
return "START $y";
sub output_record # or process it somehow...
print "Record: @record\n" if (@record >1);
Record: a b
Record: c d e f g
Record: h i
Record: j k
i START j
Posts are HTML formatted. Put <p> </p> tags around your paragraphs. Put <code> </code> tags around your code and data!
Titles consisting of a single word are discouraged, and in most cases are disallowed outright.
Read Where should I post X? if you're not absolutely sure you're posting in the right place.
Please read these before you post! —
Posts may use any of the Perl Monks Approved HTML tags:
You may need to use entities for some characters, as follows. (Exception: Within code tags, you can put the characters literally.)
- a, abbr, b, big, blockquote, br, caption, center, col, colgroup, dd, del, div, dl, dt, em, font, h1, h2, h3, h4, h5, h6, hr, i, ins, li, ol, p, pre, readmore, small, span, spoiler, strike, strong, sub, sup, table, tbody, td, tfoot, th, thead, tr, tt, u, ul, wbr
Link using PerlMonks shortcuts! What shortcuts can I use for linking?
See Writeup Formatting Tips and other pages linked from there for more info.
| & || & |
| < || < |
| > || > |
| [ || [ |
| ] || ] ||