From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from mail-pg1-x531.google.com (mail-pg1-x531.google.com [IPv6:2607:f8b0:4864:20::531]) by sourceware.org (Postfix) with ESMTPS id 4C7733858D33 for ; Mon, 25 Mar 2024 20:31:29 +0000 (GMT) DMARC-Filter: OpenDMARC Filter v1.4.2 sourceware.org 4C7733858D33 Authentication-Results: sourceware.org; dmarc=none (p=none dis=none) header.from=dabbelt.com Authentication-Results: sourceware.org; spf=pass smtp.mailfrom=dabbelt.com ARC-Filter: OpenARC Filter v1.0.0 sourceware.org 4C7733858D33 Authentication-Results: server2.sourceware.org; arc=none smtp.remote-ip=2607:f8b0:4864:20::531 ARC-Seal: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1711398691; cv=none; b=ASCujk7cTDkL1Xurg0pn+Apr0UhrqpGOKGfDb8WVH6nF3BepzWHyniltr+kfxwhyzvd4SksEP8eVtGhIK3hZ89Iune7YMjmffet30edYlnkZH/MJfKs/Or8SD3gkGrB4tQkqP98A5NTH9StGYCBjfYD2CBIGsAstS7bGyGwCXRI= ARC-Message-Signature: i=1; a=rsa-sha256; d=sourceware.org; s=key; t=1711398691; c=relaxed/simple; bh=WnIISYg+WBBlRXd/xfka3Hj1IrzdVjv2MHFsl2J4XhU=; h=DKIM-Signature:Date:Subject:From:To:Message-ID:Mime-Version; b=XLEndrxGERvtylxIhPyZLCl8TOaAIw02L9XWWvsyAUihWcwMsUZjjZ0dGM1GoAfv63D60wqoYhJxStOiBpmmZcbohzhQJumIODdlVzuM4QtmQ5bXVOw+pqCWNg1nfOxmxFCt2Dpr3h5X7Ddpj1Ja80kXNLeGtKkZPrWSaD9YK14= ARC-Authentication-Results: i=1; server2.sourceware.org Received: by mail-pg1-x531.google.com with SMTP id 41be03b00d2f7-5ce6b5e3c4eso2344804a12.2 for ; Mon, 25 Mar 2024 13:31:29 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=dabbelt-com.20230601.gappssmtp.com; s=20230601; t=1711398688; x=1712003488; darn=gcc.gnu.org; h=content-transfer-encoding:mime-version:message-id:to:from:cc :in-reply-to:subject:date:from:to:cc:subject:date:message-id :reply-to; bh=mcnKMQPSaLqpro1bzPYWZgeM7Hcf5/vvwP+kZIsDurE=; b=tUyqUlw0+QykPHZnTyYr88as7jx6TpJAgojwtm/2gqoOLDXQcdAV6FLuexjLkwRxzS ZwrvJmY7Ld0nNfhCbLgLk2UB/Ctj0BYf519OYPjJuu7Otcuyuc4s6KEjkRFj4EQxZjjL n1SpuG/i76CtSGo2StVcRhO0HMVroIxW8XZ1H0Jxm+JauY7jDdtlP+T7p0JwcL+Zsoj3 5Y0BvxJ1+XVe79Yg/gkCIm5eEXqrc3Eb2Sl1Y87nAYjje6Q8K+7vjEug18OlEXj3gn7k y20PG/+BjGER6QXzb73c7mNij8V0ERRsVwAJrXKFA8QlK6UknVNgiVxGyYHcgZm9vp8T lpeQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1711398688; x=1712003488; h=content-transfer-encoding:mime-version:message-id:to:from:cc :in-reply-to:subject:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=mcnKMQPSaLqpro1bzPYWZgeM7Hcf5/vvwP+kZIsDurE=; b=Fn+pwBojTlBI/lCGkR0Zikty83CdZcSPJHUZZ9+pNjld5axMOZZThR2MG2VdylFfbV 3lWOsoaernWpubYoF70hf8CbSk96W4RfELZGoIxykBUxdbqShY8Cxc6DSPRwADJv4Snu UCgo2TNqB/fDpB7AWJ8AdGDhwR3CwNeahDBFyKZqOVDQTMX/ui5pxDfpqqXm1CiW5q8P tqQsOZ1a+EjiCAFnodAbbwQ1aDMO3CHzqq/3KjfSeUXX3GVs4oIdB/dhK+7H8Z6/URNS irXZV3KYWPl1jKZmFCJFtcY+pv4X4oZAt9GtQBHeGHL3QUcLEDpyHo3bJX13whLr2+MG gZCg== X-Forwarded-Encrypted: i=1; AJvYcCXkaCwpznfu/gh2uUkqFc6yl4oeP5Pyv9hnRtIo+JUkH8zX2hMmijdkXUKyDLUAKU9ezukRGIUcG5HfBwN5RjyxT/is3YP2uQ== X-Gm-Message-State: AOJu0Ywyl3uWdRg1uCr+xdzxzMeFnfK+HzP+++V05DS+HO0UCb6tPdMw gjKvnxHOXWO3X9Jd/+wpQLmTrjsYTXtqdMIu/1XbsFd+brfLn8GWNk8JiFyGwBLJPdr0AmgH786 j X-Google-Smtp-Source: AGHT+IHLCQ6qQgp9gdbxexauB2dFiJEVhAU6aQmXBoOXTA+4LGJlKyhRVMAPmbhQIVo54usD6zlKWA== X-Received: by 2002:a17:902:a381:b0:1e0:a7c5:b5a5 with SMTP id x1-20020a170902a38100b001e0a7c5b5a5mr6212608pla.37.1711398687477; Mon, 25 Mar 2024 13:31:27 -0700 (PDT) Received: from localhost ([192.184.165.199]) by smtp.gmail.com with ESMTPSA id d6-20020a170902654600b001dd4d0082c8sm5241588pln.216.2024.03.25.13.31.27 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 25 Mar 2024 13:31:27 -0700 (PDT) Date: Mon, 25 Mar 2024 13:31:27 -0700 (PDT) X-Google-Original-Date: Mon, 25 Mar 2024 13:31:25 PDT (-0700) Subject: Re: TARGET_RTX_COSTS and pipeline latency vs. variable-latency instructions (was Re: [PATCH] RISC-V: Add XiangShan Nanhu microarchitecture.) In-Reply-To: <3315b6a9-96df-415d-b82f-806dada10154@ventanamicro.com> CC: xry111@xry111.site, jiawei@iscas.ac.cn, gcc-patches@gcc.gnu.org, kito.cheng@sifive.com, christoph.muellner@vrull.eu, wuwei2016@iscas.ac.cn, shihua@iscas.ac.cn, shiyulong@iscas.ac.cn, chenyixuan@iscas.ac.cn From: Palmer Dabbelt To: Jeff Law Message-ID: Mime-Version: 1.0 (MHng) Content-Type: text/plain; charset=utf-8; format=flowed Content-Transfer-Encoding: 8bit X-Spam-Status: No, score=-2.4 required=5.0 tests=BAYES_00,BODY_8BITS,DKIM_SIGNED,DKIM_VALID,RCVD_IN_DNSWL_NONE,SPF_HELO_NONE,SPF_PASS,TXREP autolearn=ham autolearn_force=no version=3.4.6 X-Spam-Checker-Version: SpamAssassin 3.4.6 (2021-04-09) on server2.sourceware.org List-Id: On Mon, 25 Mar 2024 13:27:34 PDT (-0700), Jeff Law wrote: > > > On 3/25/24 2:13 PM, Palmer Dabbelt wrote: >> On Mon, 25 Mar 2024 12:59:14 PDT (-0700), Jeff Law wrote: >>> >>> >>> On 3/25/24 1:48 PM, Xi Ruoyao wrote: >>>> On Mon, 2024-03-18 at 20:54 -0600, Jeff Law wrote: >>>>>> +/* Costs to use when optimizing for xiangshan nanhu.  */ >>>>>> +static const struct riscv_tune_param xiangshan_nanhu_tune_info = { >>>>>> +  {COSTS_N_INSNS (3), COSTS_N_INSNS (3)},    /* fp_add */ >>>>>> +  {COSTS_N_INSNS (3), COSTS_N_INSNS (3)},    /* fp_mul */ >>>>>> +  {COSTS_N_INSNS (10), COSTS_N_INSNS (20)},    /* fp_div */ >>>>>> +  {COSTS_N_INSNS (3), COSTS_N_INSNS (3)},    /* int_mul */ >>>>>> +  {COSTS_N_INSNS (6), COSTS_N_INSNS (6)},    /* int_div */ >>>>>> +  6,                        /* issue_rate */ >>>>>> +  3,                        /* branch_cost */ >>>>>> +  3,                        /* memory_cost */ >>>>>> +  3,                        /* fmv_cost */ >>>>>> +  true,                        /* slow_unaligned_access */ >>>>>> +  false,                    /* use_divmod_expansion */ >>>>>> +  RISCV_FUSE_ZEXTW | RISCV_FUSE_ZEXTH,          /* fusible_ops */ >>>>>> +  NULL,                        /* vector cost */ >>>> >>>>> Is your integer division really that fast?  The table above essentially >>>>> says that your cpu can do integer division in 6 cycles. >>>> >>>> Hmm, I just seen I've coded some even smaller value for LoongArch CPUs >>>> so forgive me for "hijacking" this thread... >>>> >>>> The problem seems integer division may spend different number of cycles >>>> for different inputs: on LoongArch LA664 I've observed 5 cycles for some >>>> inputs and 39 cycles for other inputs. >>>> >>>> So should we use the minimal value, the maximum value, or something in- >>>> between for TARGET_RTX_COSTS and pipeline descriptions? >>> Yea, early outs are relatively common in the actual hardware >>> implementation. >>> >>> The biggest reason to refine the cost of a division is so that we've got >>> a reasonably accurate cost for division by a constant -- which can often >>> be done with multiplication by reciprocal sequence.  The multiplication >>> by reciprocal sequence will use mult, add, sub, shadd insns and you need >>> a reasonable cost model for those so you can compare against the cost of >>> a hardware division. >>> >>> So to answer your question.  Choose something sensible, you probably >>> don't want the fastest case and you may not want the slowest case. >> >> Maybe we should have some sort of per-bit-set cost hook for mul/div? >> Without that we're kind of just guessing at whether the implmentation >> has early outs based on hueristics used to implicitly generate the cost >> models. >> >> Not sure that's really worth the complexity, though... > I'd doubt it's worth the complexity. Picking some reasonable value gets > you the vast majority of the benefit. Something like > COSTS_N_INSNS(6) is enough to get CSE to trigger. So what's left is a > reasonable cost, particularly for the division-by-constant case where we > need a ceiling for synth_mult. Ya, makes sense. I noticed our multi-word multiply costs are a bit odd too (they really only work for 64-bit mul on 32-bit targets), but that's probably not worth worrying about either. > > Jeff