Open
Conversation
feat(tokenizer): add support for Polish language stemming - Introduced Tantivy stemmer algorithm for Polish language alongside Rust stemmers within Language enum and stemmer logic. - Added stemmer implementation and integration for Polish. - Expanded the stop-word filter to include a list of Polish stop words. - Updated tests to validate Polish tokenizer behavior. - Included the `tantivy-stemmers` dependency to support Polish language stemming. Polish language support enhances language coverage for text processing. ``` # Conflicts: # Cargo.lock
- Eliminated an unused import for Polish language in stemmer.rs. - Cleaned up redundant code to improve maintainability.
- Applied consistent formatting to `StemmerAlgorithm` and `token_stream`. - Simplified match expressions for better readability and maintainability. - Removed unnecessary blank lines in tokenizer tests. These changes enhance code clarity and maintain coding style consistency.
Collaborator
|
Closing I am not fond of the "let's depend on two different crates" approach, especially considering it is still possible to use tantivy with polish in any program by defining your own tokenizer. I'm ok with switching to your crate though if you want to modify your PR to migrate all stemmers to yours. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This commit adds support for Polish language stemming.
Why
The previously used rust-stemmers crate is abandoned and unmaintained, which blocked the addition of new languages. This change addresses a user request for Polish stemming to improve BM25 recall in their use case. The tantivy-stemmers crate is a modern, maintained alternative that also opens the door for supporting many other languages in the future.
How
Tests