Merge pull request #36 from jntrnr/parser_improvements

Add parser README, some parser fixups
This commit is contained in:
JT 2021-09-09 07:16:24 +12:00 committed by GitHub
commit 90204bd0c8
No known key found for this signature in database
GPG key ID: 4AEE18F83AFDEB23
6 changed files with 153 additions and 19 deletions

View file

@ -17,6 +17,7 @@
- [x] Column path - [x] Column path
- [x] ...rest without calling it rest - [x] ...rest without calling it rest
- [x] Iteration (`each`) over tables - [x] Iteration (`each`) over tables
- [ ] Value serialization
- [ ] Handling rows with missing columns during a cell path - [ ] Handling rows with missing columns during a cell path
- [ ] Error shortcircuit (stopping on first error) - [ ] Error shortcircuit (stopping on first error)
- [ ] ctrl-c support - [ ] ctrl-c support

View file

@ -159,6 +159,30 @@ pub fn report_parsing_error(
Label::primary(diag_file_id, diag_range).with_message("expected type") Label::primary(diag_file_id, diag_range).with_message("expected type")
]) ])
} }
ParseError::MissingColumns(count, span) => {
let (diag_file_id, diag_range) = convert_span_to_diag(working_set, span)?;
Diagnostic::error()
.with_message("Missing columns")
.with_labels(vec![Label::primary(diag_file_id, diag_range).with_message(
format!(
"expected {} column{}",
count,
if *count == 1 { "" } else { "s" }
),
)])
}
ParseError::ExtraColumns(count, span) => {
let (diag_file_id, diag_range) = convert_span_to_diag(working_set, span)?;
Diagnostic::error()
.with_message("Extra columns")
.with_labels(vec![Label::primary(diag_file_id, diag_range).with_message(
format!(
"expected {} column{}",
count,
if *count == 1 { "" } else { "s" }
),
)])
}
ParseError::TypeMismatch(expected, found, span) => { ParseError::TypeMismatch(expected, found, span) => {
let (diag_file_id, diag_range) = convert_span_to_diag(working_set, span)?; let (diag_file_id, diag_range) = convert_span_to_diag(working_set, span)?;
Diagnostic::error() Diagnostic::error()

View file

@ -0,0 +1,99 @@
# nu-parser, the Nushell parser
Nushell's parser is a type-directed parser, meaning that the parser will use type information available during parse time to configure the parser. This allows it to handle a broader range of techniques to handle the arguments of a command.
Nushell's base language is whitespace-separated tokens with the command (Nushell's term for a function) name in the head position:
```
head1 arg1 arg2 | head2
```
## Lexing
The first job of the parser is to a lexical analysis to find where the tokens start and end in the input. This turns the above into:
```
<item: "head1">, <item: "arg1">, <item: "arg2">, <pipe>, <item: "head2">
```
At this point, the parser has little to no understanding of the shape of the command or how to parse its arguments.
## Lite parsing
As nushell is a language of pipelines, pipes form a key role in both separating commands from each other as well as denoting the flow of information between commands. The lite parse phase, as the name suggests, helps to group the lexed tokens into units.
The above tokens are converted the following during the lite parse phase:
```
Pipeline:
Command #1:
<item: "head1">, <item: "arg1">, <item: "arg2">
Command #2:
<item: "head2">
```
## Parsing
The real magic begins to happen when the parse moves on to the parsing stage. At this point, it traverses the lite parse tree and for each command makes a decision:
* If the command looks like an internal/external command literal: eg) `foo` or `/usr/bin/ls`, it parses it as an internal or external command
* Otherwise, it parses the command as part of a mathematical expression
### Types/shapes
Each command has a shape assigned to each of the arguments in reads in. These shapes help define how the parser will handle the parse.
For example, if the command is written as:
```sql
where $x > 10
```
When the parsing happens, the parser will look up the `where` command and find its Signature. The Signature states what flags are allowed and what positional arguments are allowed (both required and optional). Each argument comes with it a Shape that defines how to parse values to get that position.
In the above example, if the Signature of `where` said that it took three String values, the result would be:
```
CallInfo:
Name: `where`
Args:
Expression($x), a String
Expression(>), a String
Expression(10), a String
```
Or, the Signature could state that it takes in three positional arguments: a Variable, an Operator, and a Number, which would give:
```
CallInfo:
Name: `where`
Args:
Expression($x), a Variable
Expression(>), an Operator
Expression(10), a Number
```
Note that in this case, each would be checked at compile time to confirm that the expression has the shape requested. For example, `"foo"` would fail to parse as a Number.
Finally, some Shapes can consume more than one token. In the above, if the `where` command stated it took in a single required argument, and that the Shape of this argument was a MathExpression, then the parser would treat the remaining tokens as part of the math expression.
```
CallInfo:
Name: `where`
Args:
MathExpression:
Op: >
LHS: Expression($x)
RHS: Expression(10)
```
When the command runs, it will now be able to evaluate the whole math expression as a single step rather than doing any additional parsing to understand the relationship between the parameters.
## Making space
As some Shapes can consume multiple tokens, it's important that the parser allow for multiple Shapes to coexist as peacefully as possible.
The simplest way it does this is to ensure there is at least one token for each required parameter. If the Signature of the command says that it takes a MathExpression and a Number as two required arguments, then the parser will stop the math parser one token short. This allows the second Shape to consume the final token.
Another way that the parser makes space is to look for Keyword shapes in the Signature. A Keyword is a word that's special to this command. For example in the `if` command, `else` is a keyword. When it is found in the arguments, the parser will use it as a signpost for where to make space for each Shape. The tokens leading up to the `else` will then feed into the parts of the Signature before the `else`, and the tokens following are consumed by the `else` and the Shapes that follow.

View file

@ -28,4 +28,6 @@ pub enum ParseError {
UnknownState(String, Span), UnknownState(String, Span),
IncompleteParser(Span), IncompleteParser(Span),
RestNeedsName(Span), RestNeedsName(Span),
ExtraColumns(usize, Span),
MissingColumns(usize, Span),
} }

View file

@ -1741,13 +1741,13 @@ pub fn parse_list_expression(
pub fn parse_table_expression( pub fn parse_table_expression(
working_set: &mut StateWorkingSet, working_set: &mut StateWorkingSet,
span: Span, original_span: Span,
) -> (Expression, Option<ParseError>) { ) -> (Expression, Option<ParseError>) {
let bytes = working_set.get_span_contents(span); let bytes = working_set.get_span_contents(original_span);
let mut error = None; let mut error = None;
let mut start = span.start; let mut start = original_span.start;
let mut end = span.end; let mut end = original_span.end;
if bytes.starts_with(b"[") { if bytes.starts_with(b"[") {
start += 1; start += 1;
@ -1787,7 +1787,7 @@ pub fn parse_table_expression(
), ),
1 => { 1 => {
// List // List
parse_list_expression(working_set, span, &SyntaxShape::Any) parse_list_expression(working_set, original_span, &SyntaxShape::Any)
} }
_ => { _ => {
let mut table_headers = vec![]; let mut table_headers = vec![];
@ -1817,9 +1817,27 @@ pub fn parse_table_expression(
error = error.or(err); error = error.or(err);
if let Expression { if let Expression {
expr: Expr::List(values), expr: Expr::List(values),
span,
.. ..
} = values } = values
{ {
match values.len().cmp(&table_headers.len()) {
std::cmp::Ordering::Less => {
error = error.or_else(|| {
Some(ParseError::MissingColumns(table_headers.len(), span))
})
}
std::cmp::Ordering::Equal => {}
std::cmp::Ordering::Greater => {
error = error.or_else(|| {
Some(ParseError::ExtraColumns(
table_headers.len(),
values[table_headers.len()].span,
))
})
}
}
rows.push(values); rows.push(values);
} }
} }
@ -1828,7 +1846,7 @@ pub fn parse_table_expression(
Expression { Expression {
expr: Expr::Table(table_headers, rows), expr: Expr::Table(table_headers, rows),
span, span,
ty: Type::List(Box::new(Type::Unknown)), ty: Type::Table,
}, },
error, error,
) )
@ -2052,17 +2070,7 @@ pub fn parse_value(
} }
SyntaxShape::Any => { SyntaxShape::Any => {
if bytes.starts_with(b"[") { if bytes.starts_with(b"[") {
let shapes = [SyntaxShape::Table]; parse_value(working_set, span, &SyntaxShape::Table)
for shape in shapes.iter() {
if let (s, None) = parse_value(working_set, span, shape) {
return (s, None);
}
}
parse_value(
working_set,
span,
&SyntaxShape::List(Box::new(SyntaxShape::Any)),
)
} else { } else {
let shapes = [ let shapes = [
SyntaxShape::Int, SyntaxShape::Int,
@ -2071,8 +2079,6 @@ pub fn parse_value(
SyntaxShape::Filesize, SyntaxShape::Filesize,
SyntaxShape::Duration, SyntaxShape::Duration,
SyntaxShape::Block, SyntaxShape::Block,
SyntaxShape::Table,
SyntaxShape::List(Box::new(SyntaxShape::Any)),
SyntaxShape::String, SyntaxShape::String,
]; ];
for shape in shapes.iter() { for shape in shapes.iter() {

View file

@ -16,6 +16,7 @@ pub enum Type {
Number, Number,
Nothing, Nothing,
Record(Vec<String>, Vec<Type>), Record(Vec<String>, Vec<Type>),
Table,
ValueStream, ValueStream,
Unknown, Unknown,
Error, Error,
@ -34,6 +35,7 @@ impl Display for Type {
Type::Int => write!(f, "int"), Type::Int => write!(f, "int"),
Type::Range => write!(f, "range"), Type::Range => write!(f, "range"),
Type::Record(cols, vals) => write!(f, "record<{}, {:?}>", cols.join(", "), vals), Type::Record(cols, vals) => write!(f, "record<{}, {:?}>", cols.join(", "), vals),
Type::Table => write!(f, "table"),
Type::List(l) => write!(f, "list<{}>", l), Type::List(l) => write!(f, "list<{}>", l),
Type::Nothing => write!(f, "nothing"), Type::Nothing => write!(f, "nothing"),
Type::Number => write!(f, "number"), Type::Number => write!(f, "number"),