celeborn

Author	SHA1	Message	Date
Cheng Pan	ac84d64d51	[CELEBORN-707][MASTER] Remove env CELEBORN_MASTER_HOST and CELEBORN_MASTER_PORT ### What changes were proposed in this pull request? Remove environment variables `CELEBORN_MASTER_HOST` and `CELEBORN_MASTER_PORT`, and makes `CELEBORN_LOCAL_HOSTNAME` takes effect on both master and worker. ### Why are the changes needed? There are many different ways to configure the master/worker host and port, which makes the thing complex and inconsistent. After this change, #### master 1. cli args `--host` `--port` takes the highest priority 2. then lookup env `CELEBORN_LOCAL_HOSTNAME` 3. things are different when HA enabled and disabled 3.1. when HA is disabled, lookup configurations `celeborn.master.host` and `celeborn.master.port` 3.2. when HA is enabled, each node needs to know the whole cluster info, ``` celeborn.master.ha.node.1.host clb-1 celeborn.master.ha.node.1.port 9097 celeborn.master.ha.node.2.host clb-2 celeborn.master.ha.node.2.port 9097 celeborn.master.ha.node.3.host clb-3 celeborn.master.ha.node.3.port 9097 ``` in addition, `celeborn.master.ha.node.id=1` can be used to indicate the node id, otherwise, the master will try to bind each host to match the node id. #### worker 1. cli args `--host` `--port` takes the highest priority 2. then lookup env `CELEBORN_LOCAL_HOSTNAME` things are simple than the master case because each worker is not required to know others. ### Does this PR introduce _any_ user-facing change? Yes. ### How was this patch tested? UT. Closes #1616 from pan3793/CELEBORN-707. Authored-by: Cheng Pan <chengpan@apache.org> Signed-off-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com>	2023-06-25 16:00:59 +08:00
zky.zhoukeyong	e2eeafd4bf	[CELEBORN-709] Increase default fetch timeout ### What changes were proposed in this pull request? 30s for fetch timeout is too short and easy to exceed. This PR increases the default value to 600s. ### Why are the changes needed? When I was testing 3T TPCDS with three workers, I encountered fetch timeout: ``` 23/06/21 16:46:41,771 INFO [fetch-server-11-7] FetchHandler: Sending chunk 28856864163, 1, 0, 2147483647 ... 23/06/21 16:47:16,870 INFO [fetch-server-11-7] FetchHandler: Sent chunk 28856864163, 1, 0, 2147483647 ``` And I remember from some users' monitoring, the max fetch time can reach several minutes on heavy load without error. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Manual test. Closes #1618 from waitinfuture/709. Authored-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com> Signed-off-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com>	2023-06-23 21:06:43 +08:00
zky.zhoukeyong	5f4f6d953f	[CELEBORN-702][DOC] Extend doc about migration from 0.2.1 to 0.3.0 ### What changes were proposed in this pull request? Extend doc about migration from 0.2.1 to 0.3.0. Added the following contents: <img width="1084" alt="image" src="https://github.com/apache/incubator-celeborn/assets/26535726/7a9d172c-09ba-48b6-9f5c-73a8b13d035f"> ### Why are the changes needed? ditto ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? No. Closes #1612 from waitinfuture/702. Lead-authored-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com> Co-authored-by: Cheng Pan <pan3793@gmail.com> Signed-off-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com>	2023-06-20 20:45:58 +08:00
Cheng Pan	e22379c3ab	[CELEBORN-638] Migrate configurations celeborn.ha.master.* to celeborn.master.ha.* ### What changes were proposed in this pull request? It was discussed during the last meeting, but abandoned due to the complication. ### Why are the changes needed? Make the configuration unified. ### Does this PR introduce _any_ user-facing change? Yes, but the legacy configurations still take effect. ### How was this patch tested? New UTs. Closes #1549 from pan3793/CELEBORN-638. Authored-by: Cheng Pan <chengpan@apache.org> Signed-off-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com>	2023-06-16 18:18:26 +08:00
Angerszhuuuu	1ba6dee324	[CELEBORN-680][DOC] Refresh celeborn configurations in doc ### What changes were proposed in this pull request? Refresh celeborn configurations in doc ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1592 from AngersZhuuuu/CELEBORN-680. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Angerszhuuuu <angers.zhu@gmail.com>	2023-06-15 13:59:38 +08:00
Angerszhuuuu	0aa13832b5	[CELEBORN-676] Celeborn fetch chunk also should support check timeout ### What changes were proposed in this pull request? Celeborn fetch chunk also should support check timeout #### Test case ``` executor instance 20 SQL: SELECT count(1) from (select /+ REPARTITION(100) / * from spark_auxiliary.t50g) tmp; --conf spark.celeborn.client.spark.shuffle.writer=sort \ --conf spark.celeborn.client.fetch.excludeWorkerOnFailure.enabled=true \ --conf spark.celeborn.client.push.timeout=10s \ --conf spark.celeborn.client.push.replicate.enabled=true \ --conf spark.celeborn.client.push.revive.maxRetries=10 \ --conf spark.celeborn.client.reserveSlots.maxRetries=10 \ --conf spark.celeborn.client.registerShuffle.maxRetries=3 \ --conf spark.celeborn.client.push.blacklist.enabled=true \ --conf spark.celeborn.client.blacklistSlave.enabled=true \ --conf spark.celeborn.client.fetch.timeout=30s \ --conf spark.celeborn.client.push.data.timeout=30s \ --conf spark.celeborn.client.push.limit.inFlight.timeout=600s \ --conf spark.celeborn.client.push.maxReqsInFlight=32 \ --conf spark.celeborn.client.shuffle.compression.codec=ZSTD \ --conf spark.celeborn.rpc.askTimeout=30s \ --conf spark.celeborn.client.rpc.reserveSlots.askTimeout=30s \ --conf spark.celeborn.client.shuffle.batchHandleChangePartition.enabled=true \ --conf spark.celeborn.client.shuffle.batchHandleCommitPartition.enabled=true \ --conf spark.celeborn.client.shuffle.batchHandleReleasePartition.enabled=true ``` Test with 3 worker and add a `Thread.sleep(100s)` before worker handle `ChunkFetchRequest` Before patch <img width="1783" alt="截屏2023-06-14 上午11 20 55" src="https://github.com/apache/incubator-celeborn/assets/46485123/182dff7d-a057-4077-8368-d1552104d206"> After patch <img width="1792" alt="image" src="https://github.com/apache/incubator-celeborn/assets/46485123/3c8b7933-8ace-426d-8e9f-04e0aabfac8e"> The log shows the fetch timeout checker workers ``` 23/06/14 11:14:54 ERROR WorkerPartitionReader: Fetch chunk 0 failed. org.apache.celeborn.common.exception.CelebornIOException: FETCH_DATA_TIMEOUT at org.apache.celeborn.common.network.client.TransportResponseHandler.failExpiredFetchRequest(TransportResponseHandler.java:147) at org.apache.celeborn.common.network.client.TransportResponseHandler.lambda$new$1(TransportResponseHandler.java:103) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:748) 23/06/14 11:14:54 WARN RssInputStream: Fetch chunk failed 1/6 times for location PartitionLocation[ id-epoch:35-0 host-rpcPort-pushPort-fetchPort-replicatePort:10.169.48.203-9092-9094-9093-9095 mode:MASTER peer:(host-rpcPort-pushPort-fetchPort-replicatePort:10.169.48.202-9092-9094-9093-9095) storage hint:StorageInfo{type=HDD, mountPoint='/mnt/ssd/0', finalResult=true, filePath=} mapIdBitMap:null], change to peer org.apache.celeborn.common.exception.CelebornIOException: Fetch chunk 0 failed. at org.apache.celeborn.client.read.WorkerPartitionReader$1.onFailure(WorkerPartitionReader.java:98) at org.apache.celeborn.common.network.client.TransportResponseHandler.failExpiredFetchRequest(TransportResponseHandler.java:146) at org.apache.celeborn.common.network.client.TransportResponseHandler.lambda$new$1(TransportResponseHandler.java:103) at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:511) at java.util.concurrent.FutureTask.runAndReset(FutureTask.java:308) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.access$301(ScheduledThreadPoolExecutor.java:180) at java.util.concurrent.ScheduledThreadPoolExecutor$ScheduledFutureTask.run(ScheduledThreadPoolExecutor.java:294) at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149) at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624) at java.lang.Thread.run(Thread.java:748) Caused by: org.apache.celeborn.common.exception.CelebornIOException: FETCH_DATA_TIMEOUT at org.apache.celeborn.common.network.client.TransportResponseHandler.failExpiredFetchRequest(TransportResponseHandler.java:147) ... 8 more 23/06/14 11:14:54 INFO SortBasedShuffleWriter: Memory used 72.0 MB ``` ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1587 from AngersZhuuuu/CELEBORN-676. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Angerszhuuuu <angers.zhu@gmail.com>	2023-06-15 13:54:09 +08:00
Angerszhuuuu	8a0b7d80d6	[CELEBORN-681][DOC] Add celeborn.metrics.conf to conf entity ### What changes were proposed in this pull request? Add celeborn.metrics.conf to conf entity ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1593 from AngersZhuuuu/CELEBORN-681. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Angerszhuuuu <angers.zhu@gmail.com>	2023-06-14 18:06:03 +08:00
Fu Chen	aa3bb0ac3b	[CELEBORN-679] Optimize `Utils#bytesToString` ### What changes were proposed in this pull request? refer to https://github.com/apache/spark/pull/40301 1. Optimize `Utils.bytesToString`. Arithmetic ops on BigInt and BigDecimal are order(s) of magnitude slower than the ops on primitive types. Division is an especially slow operation and it is used en masse here. 2. According to the information sourced from [Wikipedia](https://en.wikipedia.org/wiki/Kilobyte), it is established that 1000 is the appropriate factor for representing kilobytes (KB), while 1024 is the correct factor for kibibytes (KiB). In alignment with this understanding, changing the size unit from "KB" to "KiB". ### Why are the changes needed? the Utils#bytesToString method is frequently employed in memory-related log messages. ### Does this PR introduce _any_ user-facing change? No, only perf improvement. ### How was this patch tested? existing UT and manually tested. Closes #1590 from cfmcgrady/bytesToString. Authored-by: Fu Chen <cfmcgrady@gmail.com> Signed-off-by: Cheng Pan <chengpan@apache.org>	2023-06-14 17:42:16 +08:00
Shuang	da85347330	[CELEBORN-675] Fix decode heartbeat message ### What changes were proposed in this pull request? Give Heartbeat one byte message and skip this byte when decode. ### Why are the changes needed? Heartbeat message may split in to two netty buffer, then the `empty buffer` (which don't need actually, but need keep) be wrong removed, then decodeNext would throw NPE. see ``` java while (headerBuf.readableBytes() < HEADER_SIZE) { ByteBuf next = buffers.getFirst(); int toRead = Math.min(next.readableBytes(), HEADER_SIZE - headerBuf.readableBytes()); headerBuf.writeBytes(next, toRead); if (!next.isReadable()) { buffers.removeFirst().release(); } } ``` ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? UT & MANUAL Closes #1589 from RexXiong/CELEBORN-675. Authored-by: Shuang <lvshuang.tb@gmail.com> Signed-off-by: zhongqiang.czq <zhongqiang.czq@alibaba-inc.com>	2023-06-14 14:37:13 +08:00
zky.zhoukeyong	47cded835f	[CELEBORN-669] Avoid commit files on excluded worker list ### What changes were proposed in this pull request? CommitHandler will check whether the target worker is in WorkerStatusTracker's excluded list. If so, skip calling commit files on it. ### Why are the changes needed? Avoid unnecessary commit files to excluded worker. ### Does this PR introduce _any_ user-facing change? No. ### How was this patch tested? Manual test. Closes #1581 from waitinfuture/669. Lead-authored-by: zky.zhoukeyong <zky.zhoukeyong@alibaba-inc.com> Co-authored-by: Angerszhuuuu <angers.zhu@gmail.com> Co-authored-by: Keyong Zhou <zhouky@apache.org> Signed-off-by: Shuang <lvshuang.tb@gmail.com>	2023-06-13 22:31:02 +08:00
Angerszhuuuu	357add5b00	[CELEBORN-494][PERF] RssInputStream fetch side support blacklist to avoid client side timeout in same worker multiple times during fetch ### What changes were proposed in this pull request? ####Test case ``` executor instance 20 SQL: SELECT count(1) from (select /+ REPARTITION(100) / * from spark_auxiliary.t50g) tmp; create connection timeout 10s Fetch chunk timeout 30s ``` In the graph, the shuffle read time of `before` and `after` is always the same delay time. ##### Worker can't connect Before ![image](https://user-images.githubusercontent.com/46485123/229465520-9d751b40-2b8f-49d2-b350-a2278e3dd89e.png) After ![image](https://user-images.githubusercontent.com/46485123/229465552-88ac1ca4-24ad-4c30-9a46-0cdcae6bbfd5.png) ##### OpenStream stuck Before ![image](https://user-images.githubusercontent.com/46485123/229465629-68765a6a-2503-4018-8917-d49e47d5dccc.png) After ![image](https://user-images.githubusercontent.com/46485123/229465683-2f57b374-1c66-4819-93dd-cabee7ccb788.png) ##### Fetch chunk stuck Before ![image](https://user-images.githubusercontent.com/46485123/229465735-8d2f694b-1b4a-4984-b069-c4a308f41008.png) After ![image](https://user-images.githubusercontent.com/46485123/229465754-c2237d5a-6fb6-4d5b-819e-b7d86a1e88d7.png) ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1406 from AngersZhuuuu/CELEBORN-494. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Shuang <lvshuang.tb@gmail.com>	2023-06-13 20:06:31 +08:00
Angerszhuuuu	6b725202a2	[CELEBORN-640][WORKER] DataPushQueue should not keep waiting take tasks ### What changes were proposed in this pull request? In our prod meet many times of push queue stuck caused by PushState's status was not being removed. Caused DataPushQueue to keep waiting for taking task. Although have resolved some bugs, here we'd better add a max wait time for taking tasks since we already have the `PUSH_DATA_TIMEOUT` check method. If the target worker is really stuck, we can retry our task finally. ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1552 from AngersZhuuuu/CELEBORN-640. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Angerszhuuuu <angers.zhu@gmail.com>	2023-06-09 14:06:47 +08:00
onebox-li	0c869ac9a0	[CELEBORN-642] Improve metrics and update grafana ### What changes were proposed in this pull request? Change in grafana （ALL） add: JVMCPUTime LastMinuteSystemLoad AvailableProcessors （For Master） add: LostWorkers IsActiveMaster PartitionSize （For Worker） add: PushDataFailCount -> WriteDataFailCount ReplicateDataFailCount ReplicateDataWriteFailCount ReplicateDataCreateConnectionFailCount ReplicateDataConnectionExceptionCount ReplicateDataTimeoutCount SortedFileSize PushDataHandshakeFailCount RegionStartFailCount RegionFinishFailCount MasterPushDataHandshakeTime SlavePushDataHandshakeTime MasterRegionStartTime SlaveRegionStartTime MasterRegionFinishTime SlaveRegionFinishTime PotentialConsumeSpeed UserProduceSpeed WorkerConsumeSpeed DeviceOSFreeBytes DeviceCelebornFreeBytes push usedHeapMemory/usedDirectMemory fetch usedHeapMemory/usedDirectMemory replicate usedHeapMemory/usedDirectMemory remove: dup ReserveSlotsTime Change dashboard layout. Fix support for multiple labels. Modify some metrics docs. ### Why are the changes needed? For better use of metrics. ### Does this PR introduce _any_ user-facing change? Below metrics change name, extract some value to the label. DeviceOSFreeCapacity(B) -> DeviceOSFreeBytes DeviceOSTotalCapacity(B) -> DeviceOSTotalBytes DeviceCelebornFreeCapacity(B) -> DeviceCelebornFreeBytes DeviceCelebornTotalCapacity(B) -> DeviceCelebornTotalBytes push usedHeapMemory/usedDirectMemory fetch usedHeapMemory/usedDirectMemory replicate usedHeapMemory/usedDirectMemory ### How was this patch tested? Cluster test. Closes #1557 from onebox-li/improve-metrics. Authored-by: onebox-li <lyh-36@163.com> Signed-off-by: Cheng Pan <chengpan@apache.org>	2023-06-08 18:10:06 +08:00
Ethan Feng	76a42beab0	[CELEBORN-610][FLINK] Eliminate pluginconf and merge its content to CelebornConf ### What changes were proposed in this pull request? Pluginconf might be hard to understand why Celeborn needs to config class. ### Why are the changes needed? Ditto. ### Does this PR introduce _any_ user-facing change? NO. ### How was this patch tested? UT. Closes #1524 from FMX/CELEBORN-610. Authored-by: Ethan Feng <ethanfeng@apache.org> Signed-off-by: Ethan Feng <ethanfeng@apache.org>	2023-06-05 14:08:53 +08:00
Angerszhuuuu	218bfc78a5	[CELEBORN-629][DOC] Add doc about enable rac-awareness ### What changes were proposed in this pull request? Add doc about enabling rac-awareness ### Why are the changes needed? Document new features. ### Does this PR introduce _any_ user-facing change? Yes, the docs are updated. ### How was this patch tested? <img width="1085" alt="截屏2023-06-02 下午3 19 10" src="https://github.com/apache/incubator-celeborn/assets/46485123/c8c51a4c-40be-40ea-befd-3c369b9f7600"> Closes #1536 from AngersZhuuuu/CELEBORN-629. Authored-by: Angerszhuuuu <angers.zhu@gmail.com> Signed-off-by: Cheng Pan <chengpan@apache.org>	2023-06-05 10:28:26 +08:00
Angerszhuuuu	3883fe2c80	[CELEBORN-623][FOLLUPUP] Refine doc about use ratis shell with RSS cluster ### What changes were proposed in this pull request? Refine this doc since: 1. It didn't mention our cluster default RPC type is `NETTY` 2. If the user use the ratis shell with `GRPC` but didn't know the ratis cluster is `NETTY`, the error is not clear and hard to debug. ### Why are the changes needed? ### Does this PR introduce _any_ user-facing change? ### How was this patch tested? Closes #1542 from AngersZhuuuu/CELEBORN-623-FOLLOWUP. Lead-authored-by: Angerszhuuuu <angers.zhu@gmail.com> Co-authored-by: Cheng Pan <pan3793@gmail.com> Signed-off-by: Cheng Pan <chengpan@apache.org>	2023-06-02 22:09:05 +08:00
Angerszhuuuu	4df4775524	[CELEBORN-632][DOC] Add spark name space to spark specify properties (#1538 )	2023-06-02 21:48:56 +08:00
liyihe	188b069710	[CELEBORN-623][DOCS] Document how to change RPC type in `celeborn-ratis` ### What changes were proposed in this pull request? Ratis-shell use GRPC by default. Celeborn support Netty for ratis, if `raft.rpc.type` is not specified, commands may fail. e.g. ``` org.apache.ratis.thirdparty.io.grpc.StatusRuntimeException: DEADLINE_EXCEEDED: deadline exceeded after 14.947369960s. [closed=[], open=[[buffered_nanos=14962358255, waiting_for_connection]]] ``` So I think we should update the document to mention how to change the RPC type to in `celeborn-ratis`. ### Why are the changes needed? Improve user experience ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? Manually test Closes #1530 from onebox-li/ratis-shell-default-rpc. Lead-authored-by: liyihe <liyihe@bigo.sg> Co-authored-by: Cheng Pan <pan3793@gmail.com> Signed-off-by: Cheng Pan <chengpan@apache.org>	2023-06-02 20:23:09 +08:00
Angerszhuuuu	e18a5ea769	[CELEBORN-624] StorageManager should only remove expired app dirs (#1531 )	2023-06-02 11:33:33 +08:00
Ethan Feng	d33916e571	[CELEBORN-625] Add a config to enable/disable UnsafeRow fast write. (#1532 )	2023-06-01 20:55:45 +08:00
Angerszhuuuu	cf308aa057	[CLEBORN-595] Refine code frame of CelebornConf (#1525 )	2023-06-01 10:37:58 +08:00
Angerszhuuuu	6d5dd50915	[CELEBORN-595][FOLLOWUP] Fix change version to 0.3.0. (#1522 )	2023-05-30 20:12:56 +08:00
Angerszhuuuu	62681ba85d	[CELEBORN-595] Rename and refactor the configuration doc. (#1501 )	2023-05-30 15:14:12 +08:00
zhongqiangchen	f117cff776	[CELEBORN-618] [FLINK] worker side adds partition split configuration options (#1520 )	2023-05-30 14:13:31 +08:00
Binjie Yang	d30f45ad63	[CELEBORN-450][HELM] Configurable volumes in the values.yaml (#1508 ) * [CELEBORN-450] Configure the mount & volume in the Values.yaml * fix comments * fix wrong name * fix comments * fix typo * fix into array * Wiht User Note Comments * fix comments * Update charts/celeborn/templates/worker-statefulset.yaml --------- Co-authored-by: Cheng Pan <pan3793@gmail.com>	2023-05-29 13:48:23 +08:00
Angerszhuuuu	d244f44518	[CELEBORN-593] Refine some RPC related default configurations (#1498 )	2023-05-19 18:23:12 +08:00
Angerszhuuuu	615d9a111f	[CELEBORN-487] Remove wrong space of config SHUFFLE_CLIENT_PUSH_BLACK (#1500 )	2023-05-19 14:27:57 +08:00
Angerszhuuuu	811e192bbd	[CELEBORN-446] Support rack aware during assign slots for ROUNDROBIN (#1370 )	2023-05-18 13:58:51 +08:00
Ethan Feng	7015d2463a	[CELEBORN-583] Merge pooled memory allocators. (#1490 )	2023-05-18 10:37:30 +08:00
Angerszhuuuu	791d72d45f	[CELEBORN-590] Remove hadoop prefix of WORKER_WORKING_DIR (#1494 )	2023-05-17 17:57:27 +08:00
Angerszhuuuu	7c6cb2f3bb	[CELEBORN-588] Remove test conf's category (#1491 )	2023-05-17 17:37:28 +08:00
Angerszhuuuu	64a3534f71	[CELEBORN-584] Worker side should expose push/replicate/fetch Netty allocator metrics (#1489 )	2023-05-16 17:51:33 +08:00
Angerszhuuuu	d657f8268a	[CELEBORN-586] Add SystemMiscSource to indicate system running status (#1488 )	2023-05-16 14:03:07 +08:00
zhongqiangchen	5769c3fdc7	[CELEBORN-552] Add HeartBeat between the client and worker to keep alive (#1457 )	2023-05-10 19:35:51 +08:00
Angerszhuuuu	778b5440bc	[CELEBORN-556][BUG] ReserveSlot should not use default RPC time out since register shuffle max timeout is network timeout (#1461 )	2023-05-10 12:29:06 +08:00
Ethan Feng	3e0d779962	[CELEBORN-576] Add static identity provider and manually settable identity provider for non-hadoop environment. (#1480 )	2023-05-08 17:29:01 +08:00
Ethan Feng	91b757555e	[CELEBORN-570] Update docs about monitor and deployment. (#1478 )	2023-05-08 17:07:42 +08:00
Angerszhuuuu	ef4c12e0fe	[CELEBORN-565] FETCH_MAX_RETRIES should double when enable replicates (#1471 )	2023-04-28 14:27:35 +08:00
Angerszhuuuu	13ce04f8a1	[CELEBORN-557] HA_CLIENT_RPC_ASK_TIMEOUT should fallback to RPC_ASK_TIMEOUT (#1462 ) * [CELEBORN-557] HA_CLIENT_RPC_ASK_TIMEOUT should fallback to RPC_ASK_TIMEOUT	2023-04-26 15:19:34 +08:00
Shuang	0b2e4877bd	[CELEBORN-553] Improve IO (#1458 )	2023-04-25 21:14:06 +08:00
Angerszhuuuu	0c2d3e647d	[CELEBORN-532][METRICS] Refine push-related failure metrics (#1442 ) * [CELEBORN-532][METRICS] Refine push-related failure metrics	2023-04-21 17:05:43 +08:00
Angerszhuuuu	181c1bfcd6	[CELEBORN-524][PERF] CongestionControl call too much ChannelsLimiter onTrim cause CPU stuck or occupy too much CPU cause no cpu for handlePushData (#1428 )	2023-04-21 15:44:56 +08:00
Angerszhuuuu	6830cb61ef	[CELEBORN-540][Refactor] Add config entity of celeborn.rpc.io.threads (#1443 ) * [CELEBORN-540][CONF] Add config entity of celeborn.rpc.io.threads	2023-04-21 11:21:41 +08:00
Angerszhuuuu	e319b99a1c	[CELEBORN-527][DOC] Fix incorrect monitor the arrangement of documents (#1432 )	2023-04-17 11:12:19 +08:00
Angerszhuuuu	ecafbf41fc	[CELEBORN-516][FOLLOWUP] Remove removed RPC metrics in metric doc (#1431 )	2023-04-17 10:46:04 +08:00
cxzl25	13f772e0c0	[CELEBORN-525] Fix wrong parameter celeborn.push.buffer.size	2023-04-14 20:45:25 +08:00
Cheng Pan	fb7b311c89	[CELEBORN-499] Move version specific resource to main repo (#1429 ) * [CELEBORN-499] Move version specific resource to main repo * license	2023-04-14 16:20:51 +08:00
Ethan Feng	9cccfc9872	[CELEBORN-431][FLINK] Support dynamic buffer allocation in reading map partition. (#1407 )	2023-04-13 10:37:47 +08:00
Angerszhuuuu	e5722126e9	[CELEBORN-502][REFACTOR] Merge GetBlacklistResponse to HeartbeatFromApplication (#1408 ) * [CELEBORN-502][REFACTOR] Merge GetBlacklistResponse to HeartbeatFromApplication	2023-04-12 14:59:32 +08:00
Angerszhuuuu	cad2836e85	[CELEBORN-505] Fix typo of SHUFFLE_CHUCK_SIZE (#1411 )	2023-04-04 19:15:30 +08:00

1 2 3

130 Commits