elasticsearch

mirror of https://github.com/elastic/elasticsearch.git synced 2025-04-25 15:47:23 -04:00

Author	SHA1	Message	Date
debadair	777598d602	[DOCS] Remove redirect pages (#88738 ) * [DOCS] Remove manual redirects * [DOCS] Removed refs to modules-discovery-hosts-providers * [DOCS] Fixed broken internal refs * Fixing bad cross links in ES book, and adding redirects.asciidoc[] back into docs/reference/index.asciidoc. * Update docs/reference/search/point-in-time-api.asciidoc Co-authored-by: James Rodewig <james.rodewig@elastic.co> * Update docs/reference/setup/restart-cluster.asciidoc Co-authored-by: James Rodewig <james.rodewig@elastic.co> * Update docs/reference/sql/endpoints/translate.asciidoc Co-authored-by: James Rodewig <james.rodewig@elastic.co> * Update docs/reference/snapshot-restore/restore-snapshot.asciidoc Co-authored-by: James Rodewig <james.rodewig@elastic.co> * Update repository-azure.asciidoc * Update node-tool.asciidoc * Update repository-azure.asciidoc --------- Co-authored-by: amyjtechwriter <61687663+amyjtechwriter@users.noreply.github.com> Co-authored-by: Elastic Machine <elasticmachine@users.noreply.github.com> Co-authored-by: Amy Jonsson <amy.jonsson@elastic.co> Co-authored-by: James Rodewig <james.rodewig@elastic.co>	2023-05-24 12:32:46 +01:00
Kostas Krikellas	deffa800db	Support value retrieval in top_hits (#95828 ) This is used when the `top_hits` output is passed to pipeline aggregators like bucket selectors. The logic retrieves the requested field from the source of the first SearchHit. This implies that (a) the spec of the wrapping aggregator (e.g. `bucket_path`) points to an appropriate field using a bracketed reference (e.g. `my_top_hits[my_metric]`) and (b) the `top_hits` contains a `size: 1` setting. This PR also includes extensions to YAML tests for `top_metrics` and `top_hits` to cover the cases where these are used in pipeline aggregations through `bucket_selector`, similar to a HAVING clause in SQL. Related to https://github.com/elastic/elasticsearch/issues/73429.	2023-05-15 09:21:11 -04:00
tmgordeeva	2abbce0e50	Time series docs (#94337 ) * Time series docs Tech preview docs with a very basic example. --------- Co-authored-by: lcawl <lcawley@elastic.co>	2023-05-03 11:01:07 -07:00
QY	2306f78ca9	Add `keyed` param to allow named filters agg return buckets as an array of objects (#89256 ) Adds a new `keyed` param for `filters` aggs to come back with their `key` attached rather than as a json object. So that sorting them is meaningful.	2023-04-10 13:53:26 -04:00
István Zoltán Szabó	b0a275dee6	[DOCS] Adds tip to change point agg docs. (#94981 )	2023-04-05 15:43:16 +02:00
István Zoltán Szabó	0405a4d1f0	[DOCS] Adds an example to change point aggregation (#94776 ) * [DOCS] Adds an example to change point aggregation. * [DOCS] Edits.	2023-03-27 18:06:11 +02:00
iamthinh	7559630e65	Update pipeline.asciidoc (#94526 ) Remove redundant character	2023-03-21 09:47:55 +01:00
Craig Taverner	f55d70a682	Document datehistogram with long offsets (#93328 ) * Document datehistogram with long offsets When offsets are longer than calendar_intervals that are non-standard, like months which differ in length, then the usual rule of all buckets starting at the same day and time will no longer apply. This update attempts to explain this with examples. * Removed TEST-skip lines These don't seem to be parsable, even though they match the syntax described in the README.asciidoc * Added // TESTRESPONSE[skip:...] lines * Refined docs description and added more examples * Update docs/reference/aggregations/bucket/datehistogram-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co> * Update docs/reference/aggregations/bucket/datehistogram-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co> * Update docs/reference/aggregations/bucket/datehistogram-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co> * Update docs/reference/aggregations/bucket/datehistogram-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co> --------- Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co>	2023-02-06 16:20:40 +01:00
Hendrik Muhs	cf5ea0bb1f	[ML] rename frequent_items to frequent_item_sets and make it GA (#93421 ) rename frequent_items to frequent_item_sets and remove the experimental batch	2023-02-02 09:25:00 +01:00
Glen Smith	81d9cbe0ca	Update frequent-items-aggregation.asciidoc (#93287 ) Fix type togeher > together	2023-01-27 09:45:17 -05:00
Craig Taverner	e8b4de9a8a	Documentation for geohex_grid over geo_shape (#92999 ) * Documentation for geohex_grid over geo_shape The feature to add support for geohex_grid aggregations over geo_shape fields was added in https://github.com/elastic/elasticsearch/pull/91956. This is the associated documentation for that. * Update docs/reference/aggregations/bucket/geohexgrid-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co> * Fix explanation for geo_point vs geo_shape proj When aggregating geohex over geoshape we use requirectangular because underlying lucene index indexes and searches the polygons in that way. * Correct spelling According to grammarly, "therefor" is not an alternative spelling of "therefore". We should use the conjunctive form here. See https://www.grammarly.com/blog/therefore-vs-therefor/ Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co>	2023-01-24 16:03:27 +01:00
István Zoltán Szabó	e4721f1dfe	[DOCS] Fine-tunes documentation on exclude/include in frequent items (#92758 )	2023-01-10 12:23:27 +01:00
Hendrik Muhs	b9c0315d24	[ML] add the ability to include and exclude values in Frequent items (#92414 ) This PR adds include and excludes to frequent items. This will allow to filter values from the analysis.	2022-12-21 12:24:10 +01:00
Paweł Krześniak	34c30ad7be	[DOCS] typo in date_histogram aggregation example (#91715 ) * [DOCS] typo in date_histogram aggregation example The field name fixed * Update docs/reference/aggregations/bucket/datehistogram-aggregation.asciidoc Co-authored-by: Abdon Pijpelink <abdon.pijpelink@elastic.co>	2022-11-21 13:13:44 +01:00
Craig Taverner	81d5859f61	Added documentation for cartesian-bounds aggregation (#91623 ) * Added documentation for cartesian-bounds aggregation * Fixed rounding errors in docs tests	2022-11-18 11:00:41 +01:00
Lisa Cawley	d7c0b37924	[DOCS] Edits frequent items aggregation (#91564 )	2022-11-14 17:20:27 -08:00
David Roberts	3dbaa3ff23	[ML] Make categorize_text aggregation GA (#88600 ) Removes the experimental tag from the categorize_text aggregation.	2022-11-09 13:05:35 +00:00
Hendrik Muhs	14b2d2d37e	[ML] frequent items filter (#91137 ) add a filter to the frequent items agg that filters documents from the analysis while still calculating support on the full set A filter is specified top-level in frequent_items: "frequent_items": { "filter": { "term": { "host.name.keyword": "i-12345" } }, ... The above filters documents that don't match, however still counts the docs when calculating support. That's in contrast to specifying a query at the top, in which case you find the same item sets, but don't know the importance given the full document set.	2022-11-03 13:58:40 +01:00
David Roberts	be006e2eee	[ML] Improve categorize_text docs (#90765 ) Adds more detail about the meaning of the results fields of the `categorize_text` aggregation, and advice about how to use these fields when searching for messages that match the categories. Followup to #90723	2022-10-13 10:46:53 +01:00
David Roberts	bfccd20155	[ML] Add a regex to the output of the categorize_text aggregation (#90723 ) The new `regex` field in `categorize_text` output is created in the same way as the `regex` field that appears in the category definitions created by anomaly detection jobs that do categorization. It consists of the terms that occur in the same order for every message that matches the category, separated with a `.+?` wildcard. It therefore matches the category messages and enforces the order of the terms that occurred in the same order for all messages used to create the category. It is not recommended to use the regex as the primary mechanism for searching for the original documents that were categorized. Search using a regular expression is very slow. Instead the terms of the category should be used to search for matching documents, as a terms search can use the inverted index and hence be much faster. However, there may be situations where it is useful to use the `regex` field to test whether a small set of messages that have not been indexed match the category.	2022-10-10 11:41:16 +01:00
Craig Taverner	4c5d24610f	Centroid aggregation for cartesian points and shapes (#89216 ) Added Cartesian support for centroid aggregation * First draft of cartesian-centroid docs However, this is largely a duplicate of geo-centroid docs since they are essentially identical behaviour. We should consider merging them. * Work on isAggregatable caused a minor logic conflict. When that work was done, Point and Shape were not aggregatable, but now they are.	2022-09-28 17:14:30 +02:00
Nik Everett	0683c90ded	REST tests for normalize agg (#89629 ) This adds a REST test for the normalize pipeline agg so we have backwards compatibility tests for it.	2022-08-26 14:18:46 -04:00
István Zoltán Szabó	7602015384	[DOCS] Improves frequent items aggregation docs (#89122 )	2022-08-08 15:46:29 +02:00
Benjamin Trent	46fc42b817	[ML] Make bucket_count_ks_test aggregation generally available (#88657 ) Initially released in 7.14, bucket_count_ks_test is now generally available.	2022-07-25 13:30:48 -04:00
Benjamin Trent	239d45a019	[ML] make bucket_correlation aggregation generally available (#88655 ) Originally released in 7.14, bucket_correlation is now generally available.	2022-07-21 07:20:09 -04:00
Benjamin Trent	94f2544998	Adding cardinality support for random_sampler agg (#86838 ) This adds support for the `cardinality` aggregation within a random_sampler. This usecase is helpful in determining the ratio of unique values compared to the count of total documents within the sampled set.	2022-07-21 07:19:35 -04:00
Sean Letendre	67cacde18b	Corrected an incomplete sentence. (#86542 ) * Corrected an incomplete sentence. * Update docs/reference/aggregations/metrics/avg-aggregation.asciidoc Co-authored-by: Christos Soulios <1561376+csoulios@users.noreply.github.com> Co-authored-by: David Kilfoyle <41695641+kilfoyle@users.noreply.github.com> Co-authored-by: Christos Soulios <1561376+csoulios@users.noreply.github.com>	2022-07-12 09:19:58 -04:00
Mark Tozzi	9ee6a19187	Add ability to select execution mode for cardinality aggregation (#87704 ) Plumbs through a new parameter for the cardinality aggregation, to allow configuring the execution mode. This can have significant impacts on speed and memory usage. This PR exposes three collection modes and two heuristics that we can tune going forward. All of these are treated as hints and can be silently ignored, e.g. if not applicable to the given field type. I've change the default behavior to optimize for time, which potentially uses more memory. Users can override this for the old behavior if needed.	2022-07-05 09:11:22 -04:00
apeltop	71234f7464	[DOCS] Fix typos in docs (#88226 )	2022-07-05 11:02:29 +02:00
David Roberts	93bc2e382f	[ML] Replace the implementation of the categorize_text aggregation (#85872 ) This replaces the implementation of the categorize_text aggregation with the new algorithm that was added in #80867. The new algorithm works in the same way as the ML C++ code used for categorization jobs (and now includes the fixes of elastic/ml-cpp#2277). The docs are updated to reflect the workings of the new implementation.	2022-05-23 18:46:13 +01:00
Umut Uz	53461f89f1	Remove duplicate text from cardinality aggs docs (#86615 ) The same explanation is repeated twice within a section.	2022-05-19 11:51:31 -07:00
Craig Taverner	5f7ea792ac	Soft-deprecation of point/geo_point formats (#86835 ) * Soft-deprecation of point/geo_point formats Since GeoJSON and WKT are now common formats for all three types: geo_shape, geo_point and point We decided to soft-deprecate the other point formats by ordering: * GeoJSON (object with keys `type` and `coordinates`) * WKT `POINT(x y)` * Object with keys `lat` and `lon` (or `x` and `y` for point) * Array [lon,lat] * String `"lat,lon"` (or `"x,y"` in point) * String with geohash (only in `geo_point`) The geohash is last because it is only in one field type. The string version is second last because it is the most controversial being the only version to reverse the coordinate order from all other formats (for geo_point only, since the coordinates are not reversed in point). In addition we replaced many examples in both documentation and tests to prioritize WKT over the plain string format. Many remaining examples of array format or object with keys still exist and could be replaced by, for example, GeoJSON, if we feel the need. * Incorrect quote position	2022-05-17 23:46:43 +02:00
Mark Tozzi	54efc59eff	Clarify risks around ordering terms aggregation (#86528 ) Add some details as to why some terms orderings are worse than others. Co-authored-by: Adam Locke <adam.locke@elastic.co>	2022-05-16 11:05:22 -04:00
István Zoltán Szabó	95ef40656f	[DOCS] Adds more details to the frequent items agg documentation (#86661 ) Co-authored-by: Mark Tozzi <mark.tozzi@gmail.com>	2022-05-16 10:24:14 +02:00
István Zoltán Szabó	e590e900a4	[DOCS] Adds frequent items agg docs (#86037 ) Co-authored-by: Lisa Cawley <lcawley@elastic.co>	2022-05-05 16:07:24 +02:00
Benjamin Trent	237e345d71	[ML][Docs] fix minimum buckets for change_point agg (#86396 )	2022-05-04 09:37:46 -04:00
Benjamin Trent	c49b92e425	Allow bucket paths to specify _count within a bucket (#85720 ) Users should be able to specify specific metrics/keys within a specific bucket key. An example is `agg["bucket_foo"]._count`. This change now allows that. closes: https://github.com/elastic/elasticsearch/issues/76320	2022-04-29 08:42:46 -04:00
James Garside	fca3487395	Updated format parameter description to reference Java decimal format (#86163 )	2022-04-25 20:52:44 +01:00
Elasticsearch addict	7b2511e22b	Update histogram-aggregation.asciidoc (#85356 ) Fix small grammatical mistake. Closes #85355	2022-03-28 12:27:32 -07:00
Salvatore Campagna	db6c58ed45	fix: use the correct field name when reading data from multi fields (#84752 ) When using a multi-field we need to extract data from the document using the correct field name. That is the name of the top field. Here we delegate extraction of the correct name to a method in the SearchContext that is wrapped by the AggregationContext. Issue: #82918	2022-03-11 17:11:26 +01:00
Abele Mălan	9ecb96fcf3	Fix some typos in plugins & reference docs (#84667 ) This pull request removes a few instances of duplicate words or punctuation and erroneous spelling from the docs.	2022-03-07 12:29:58 -05:00
Benjamin Trent	cf151b53fe	[ML] adds new change_point pipeline aggregation (#83428 ) adds a new `change_point` sibling pipeline aggregation. This aggregation detects a change_point in a multi-bucket aggregation. Example: ``` POST kibana_sample_data_flights/_search { "size": 0, "aggs": { "histo": { "date_histogram": { "field": "timestamp", "fixed_interval": "3h" }, "aggs": { "ticket_price": { "max": { "field": "AvgTicketPrice" } } } }, "changes": { "change_point": { "buckets_path": "histo>ticket_price" } } } } ``` Response ``` { /<snip>/ "aggregations" : { "histo" : { "buckets" : [ /<snip>/ ] }, "changes" : { "bucket" : { "key" : "2022-01-28T23:00:00.000Z", "doc_count" : 48, "ticket_price" : { "value" : 1187.61083984375 } }, "type" : { "distribution_change" : { "p_value" : 0.023753965139433175, "change_point" : 40 } } } } } ```	2022-03-04 07:00:58 -05:00
Benjamin Trent	b592d2bf01	New random_sampler aggregation for sampling documents in aggregations (#84363 ) This adds a new sampling aggregation that performs a background sampling over all documents in an index. The syntax is as follows: ``` { "aggregations": { "sampling": { "random_sampler": { "probability": 0.1 }, "aggs": { "price_percentiles": { "percentiles": { "field": "taxful_total_price" } } } } } } ``` This aggregation provides fast random sampling over the entire document set in order to speed up costly aggregations. Testing this over a variety of aggregations and data sets, the median speed up when sampling at `0.001` over millions of documents is around 70X speed improvement. Relative error rate does rely on the size of the data and the aggregation kind. Here are some typically expected numbers when sampling over 10s of millions of documents. `p` is the configured probability and `n` is the number of documents matched by your provided filter query.	2022-03-02 14:32:30 -05:00
James Rodewig	74e4add3a8	[DOCS] Update sum aggregation for histograms (#84493 ) (#84496 ) Fixes an error and test snippets for the sum aggregation example for histograms. Closes #84491 Co-authored-by: James Rodewig <40268737+jrodewig@users.noreply.github.com> (cherry picked from commit `fb45ac9dea`) Co-authored-by: Maja Grubic <maja.grubic@elastic.co>	2022-03-01 08:42:05 -05:00
Lisa Cawley	4fbbcda494	[DOCS] Fix nesting in bucket correlation aggregation (#83816 )	2022-02-11 11:14:11 -08:00
James Rodewig	d31bdd6bf4	[DOCS] Remove unneeded callouts from snippets (#83798 ) These callouts aren't referenced anywhere. Leaving them in can be confusing.	2022-02-10 15:04:46 -05:00
James Rodewig	280fd2fff7	[DOCS] Fix min/max agg snippets for histograms (#83695 ) * Updates the `min` and `max` snippets for histograms. These should now run as docs integration tests. * Fixes a copy/paste error in the `max` aggregation snippet for histograms. Relates to https://github.com/elastic/elasticsearch/pull/83384	2022-02-08 19:48:15 -05:00
Salvatore Campagna	9de75c2ac5	Add an aggregator for IPv4 and IPv6 subnets (#82410 ) Parameters accepted by the aggregator include: * prefix_length (integer, required): defines the network size of the subnet mask; * is_ipv6 (boolean, optional, default: false): defines whether the prefix applies to IPv6 (true) or IPv4 (false) IP addresses; * min_doc_count (integer, optional, default: 1): defines the minimum number of documents for a bucket to be returned in the results; * append_prefix_length (boolean, optional, default: false): defines if the prefix length is appended to the IP address key when returning results; * keyed (boolean, optional, default: false): defines whether the result is returned keyed or as an array of buckets; Each bucket returned by the aggregator represents a different subnet. IPv4 subnets also include a netmask field set to the subnet mask value (i.e. "255.255.0.0" for a /16 subnet). Related to: #57964 and elastic/kibana#68424	2022-01-28 11:59:07 +01:00
Ignacio Vera	0873893bb7	New GeoHexGrid aggregation (#82924 ) This commit introduces a new geogrid aggregation called GeoHexGridAggregation that is based in Uber h3 grid. It only supports geo_point fields.	2022-01-27 07:45:51 +01:00
James Rodewig	63f228e24e	[DOCS] Re-add paragraph noting `doc_count` is approximate (#83154 ) This paragraph was accidentally removed as part of #79205. Also fixes a minor heading capitalization error.	2022-01-26 11:07:59 -05:00

1 2 3 4 5 ...

553 commits